Full transcript
0:00Stop buying overpriced graphics cards
0:02right now. You might think you need a
0:04massive expensive video card to run
0:06local artificial intelligence models.
0:08Everyone tells you that giant hardware
0:10is the only way to get real performance
0:12on your desks. They are completely
0:14wrong. Three specific technological
0:17shifts are about to make local
0:18intelligence astonishingly cheap. The
0:21hardware monopoly is finally breaking.
0:23The path forward looks completely
0:25different from what you expect.
0:27Big tech companies are spending up to
0:30$800 billion on artificial intelligence
0:32infrastructure. Most of this capital
0:35buys enterprise chips like the Nvidia
0:37H100. For a long time, the real
0:39production bottleneck was advanced chip
0:41packaging. Building each accelerator
0:43required stacking high bandwidth memory
0:45on special substrates. Today, packaging
0:48capacity has expanded to 140,000 wafers
0:51per month. This surge in supply caused
0:54cloud rental rates to plummet. Renting
0:56an H100 dropped from $10 per hour to
0:58under $4. Many people think this drop
1:01means cheap graphics cards will soon
1:03flood the used market. They assume
1:05corporate hardware will end up in
1:06desktop computers. That will not happen
1:09because modern corporate accelerators do
1:11not use standard desktop connectors.
1:13They come in proprietary server form
1:15factors. You cannot plug an H100 into a
1:18normal consumer motherboard. Even when
1:20companies retire this equipment, it will
1:22remain trapped inside remote data
1:24centers. Some builders try to buy
1:26ancient enterprise scrap instead. You
1:28can find a Tesla P40 from 2016 for $250.
1:32It has 24 GB of memory but lacks modern
1:35tensor cores. Generating text on a P40
1:38slows down to three words per second.
1:40This makes old enterprise hardware
1:42practically useless for fast interactive
1:44tasks. Another massive barrier keeps
1:46consumer graphics card prices
1:48artificially high today.
1:50The biggest obstacle to affordable
1:52hardware is the semiconductor memory
1:54market. We are in the middle of a
1:56massive memory super cycle. Modern
1:59architectures require enormous amounts
2:01of high bandwidth memory. Memory
2:03manufacturers shifted their fabrication
2:05lines to chase lucrative server
2:07contracts. Data centers now consume
2:09about 70% of the entire global memory
2:12output. This created a severe shortage
2:14of standard memory modules used on
2:16consumer cards. A single new generation
2:19memory chip now costs up to $70. That is
2:22more than triple the $20 price of
2:25previous generations. Graphics card
2:27makers pass this extreme production cost
2:29directly to you. To make matters worse,
2:31the three dominant memory suppliers face
2:34serious legal challenges. Major
2:36companies were hit with a class action
2:38lawsuit accusing them of price fixing.
2:40The plaintiffs argue these companies
2:42intentionally cut regular memory
2:44production to inflate market prices.
2:46Analysts expect this price squeeze to
2:48continue until the end of 2027.
2:51If you find this hardware breakdown
2:53helpful, please subscribe and tap the
2:55like button. Your support helps me track
2:57these market shifts. With retail prices
3:00remaining terrible builders must look at
3:02used options. The premier survival
3:04choice is the Nvidia RTX3090
3:07released 6 years ago. It provides 24 GB
3:10of fast memory. You can buy three used
3:123090 cards for the price of one used
3:144090. that gives you 72 GB of video
3:17memory to run massive models, but
3:19stacking used graphics cards creates a
3:21massive new problem you have to solve.
3:24A single RTX 3090 pulls 350 W under
3:29heavy load. A dual card workstation
3:31easily draws 1 kW of power from the
3:34wall. In places with high utility rates,
3:36electricity expenses add up rapidly.
3:40Running a continuous 500 W machine adds
3:42massive costs to your monthly power
3:44bill. Power instability might even force
3:47you to buy expensive battery backup
3:49stations. Fortunately, modern software
3:51optimization offers a clever escape from
3:54this hardware trap. The first
3:56breakthrough lowering the barrier is
3:57dynamic expert offloading. Old models
4:00activated every single parameter in the
4:02network for every calculation. Modern
4:05designs use a mixture of experts
4:07architecture instead. Leading models
4:09contain nearly 700 billion parameters in
4:11total. Yet they only activate 37 billion
4:14parameters to produce each individual
4:16token. The computing workload is light,
4:19but storing all those parameters
4:21requires massive memory. Inference
4:23engines solve this puzzle through
4:25intelligent memory division. They place
4:27the most active experts inside your fast
4:29graphics memory. Inactive experts stay
4:32inside your regular inexpensive system
4:34memory. As the model generates text, the
4:37required experts stream across the
4:39motherboard on demand. This introduces a
4:41tiny delay, but makes massive models
4:43runnable on normal computers. Updates to
4:46AMD software also delivered a 60%
4:48performance boost on Radiant cards. This
4:51makes non- Nvidia graphics cards viable
4:53options for local workflows, but
4:55shuttling data across a motherboard bus
4:57still creates a physical bottleneck.
5:00The second major technology changing
5:02this market is unified memory
5:04architecture. In traditional computers,
5:06the central processor and graphics card
5:08maintain separate memory pools. Data
5:11must constantly travel between them over
5:13a constrained connection. Unified memory
5:16eliminates that barrier by placing both
5:18processors on a single memory pool.
5:20Apple proved this concept at the high
5:22end with their custom silicon. A Mac
5:24Studio supports up to 512 GB of memory.
5:28It provides incredible bandwidth
5:29exceeding 1 TBTE per second. That
5:32machine runs giant models completely
5:34inside local memory without quality
5:36loss. But spending $5,000 on Apple
5:38hardware is impossible for most
5:40builders. The real consumer breakthrough
5:42comes from AMD with their upcoming
5:44stricks Halo platform. This hybrid
5:47processor packs 16 high performance
5:49cores alongside a powerful graphics
5:51unit. The chip interfaces directly with
5:53up to 128 GB of fast memory. Because the
5:57memory is unified, you do not need
5:58dedicated video memory. The processor
6:01communicates at 256 GB per second across
6:04a wide interface. [music] This unified
6:06layout completely bypasses the
6:08traditional bottlenecks that plague
6:09regular desktop computers. You can go
6:12into the system settings and dedicate 96
6:14GB entirely to graphics. This allows a
6:17tiny energyefficient computer to run a
6:1970 billion parameter model. Perfectly.
6:22Best of all, the entire platform
6:23operates at roughly 120 W. You no longer
6:26need to purchase noisy industrial power
6:28units or multipleuse graphics cards, but
6:31an extraordinary mathematical
6:32breakthrough attacks the problem from
6:34the algorithmic side.
6:36The third and most radical development
6:38is the arrival of one bit language
6:40models. Researchers recently created a
6:43design known as bitnet B158.
6:46Traditional artificial intelligence
6:47models store weights as complex
6:49floatingoint numbers. Calculating those
6:52values requires massive matrix
6:54multiplication that only specialized
6:56processors can handle. Bitnet changes
6:58this fundamental math by training models
7:00with simple turnary values. Every single
7:03weight in the network is restricted to
7:05negative 1 0 or positive 1. The system
7:08entirely eliminates expensive floating
7:10point multiplication. When a weight is
7:12positive 1, the processor simply adds
7:14the incoming value. When a weight is
7:16negative one, the processor subtracts
7:18it. When a weight is zero, the
7:20calculation is skipped completely. This
7:22simple arithmetic cuts the energy cost
7:25of matrix operations by more than 70
7:27times. Memory consumption drops by up to
7:3090%. A 2 billion parameter model that
7:33normally takes 6 GB shrinks down to 400
7:36MGB. Testers successfully ran this model
7:39on a cheap single board computer at 40
7:42tokens per second. Recent research
7:44confirms this architecture scales
7:46successfully up to massive model sizes.
7:49If models no longer require floatingoint
7:51multiplication, expensive graphics cards
7:53lose their primary purpose. By 2028, the
7:56semiconductor memory shortage will
7:58finally end. That price reduction will
8:00coincide with production ready 1-bit
8:02models running on standard processors.
8:05You will no longer need power- hungry
8:07graphics cards to experience top tier
8:09intelligence locally. Running a massive
8:11model on your personal computer will
8:13soon become incredibly simple. Which
8:15large language model are you most
8:17excited to run locally once hardware
8:19barriers disappear?