Free YouTube Transcribe

Video transcript

Three technologies that will finally make local AI cheap

DevLogic · 1,260 words · 6 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00Stop buying overpriced graphics cards

0:02right now. You might think you need a

0:04massive expensive video card to run

0:06local artificial intelligence models.

0:08Everyone tells you that giant hardware

0:10is the only way to get real performance

0:12on your desks. They are completely

0:14wrong. Three specific technological

0:17shifts are about to make local

0:18intelligence astonishingly cheap. The

0:21hardware monopoly is finally breaking.

0:23The path forward looks completely

0:25different from what you expect.

0:27Big tech companies are spending up to

0:30$800 billion on artificial intelligence

0:32infrastructure. Most of this capital

0:35buys enterprise chips like the Nvidia

0:37H100. For a long time, the real

0:39production bottleneck was advanced chip

0:41packaging. Building each accelerator

0:43required stacking high bandwidth memory

0:45on special substrates. Today, packaging

0:48capacity has expanded to 140,000 wafers

0:51per month. This surge in supply caused

0:54cloud rental rates to plummet. Renting

0:56an H100 dropped from $10 per hour to

0:58under $4. Many people think this drop

1:01means cheap graphics cards will soon

1:03flood the used market. They assume

1:05corporate hardware will end up in

1:06desktop computers. That will not happen

1:09because modern corporate accelerators do

1:11not use standard desktop connectors.

1:13They come in proprietary server form

1:15factors. You cannot plug an H100 into a

1:18normal consumer motherboard. Even when

1:20companies retire this equipment, it will

1:22remain trapped inside remote data

1:24centers. Some builders try to buy

1:26ancient enterprise scrap instead. You

1:28can find a Tesla P40 from 2016 for $250.

1:32It has 24 GB of memory but lacks modern

1:35tensor cores. Generating text on a P40

1:38slows down to three words per second.

1:40This makes old enterprise hardware

1:42practically useless for fast interactive

1:44tasks. Another massive barrier keeps

1:46consumer graphics card prices

1:48artificially high today.

1:50The biggest obstacle to affordable

1:52hardware is the semiconductor memory

1:54market. We are in the middle of a

1:56massive memory super cycle. Modern

1:59architectures require enormous amounts

2:01of high bandwidth memory. Memory

2:03manufacturers shifted their fabrication

2:05lines to chase lucrative server

2:07contracts. Data centers now consume

2:09about 70% of the entire global memory

2:12output. This created a severe shortage

2:14of standard memory modules used on

2:16consumer cards. A single new generation

2:19memory chip now costs up to $70. That is

2:22more than triple the $20 price of

2:25previous generations. Graphics card

2:27makers pass this extreme production cost

2:29directly to you. To make matters worse,

2:31the three dominant memory suppliers face

2:34serious legal challenges. Major

2:36companies were hit with a class action

2:38lawsuit accusing them of price fixing.

2:40The plaintiffs argue these companies

2:42intentionally cut regular memory

2:44production to inflate market prices.

2:46Analysts expect this price squeeze to

2:48continue until the end of 2027.

2:51If you find this hardware breakdown

2:53helpful, please subscribe and tap the

2:55like button. Your support helps me track

2:57these market shifts. With retail prices

3:00remaining terrible builders must look at

3:02used options. The premier survival

3:04choice is the Nvidia RTX3090

3:07released 6 years ago. It provides 24 GB

3:10of fast memory. You can buy three used

3:123090 cards for the price of one used

3:144090. that gives you 72 GB of video

3:17memory to run massive models, but

3:19stacking used graphics cards creates a

3:21massive new problem you have to solve.

3:24A single RTX 3090 pulls 350 W under

3:29heavy load. A dual card workstation

3:31easily draws 1 kW of power from the

3:34wall. In places with high utility rates,

3:36electricity expenses add up rapidly.

3:40Running a continuous 500 W machine adds

3:42massive costs to your monthly power

3:44bill. Power instability might even force

3:47you to buy expensive battery backup

3:49stations. Fortunately, modern software

3:51optimization offers a clever escape from

3:54this hardware trap. The first

3:56breakthrough lowering the barrier is

3:57dynamic expert offloading. Old models

4:00activated every single parameter in the

4:02network for every calculation. Modern

4:05designs use a mixture of experts

4:07architecture instead. Leading models

4:09contain nearly 700 billion parameters in

4:11total. Yet they only activate 37 billion

4:14parameters to produce each individual

4:16token. The computing workload is light,

4:19but storing all those parameters

4:21requires massive memory. Inference

4:23engines solve this puzzle through

4:25intelligent memory division. They place

4:27the most active experts inside your fast

4:29graphics memory. Inactive experts stay

4:32inside your regular inexpensive system

4:34memory. As the model generates text, the

4:37required experts stream across the

4:39motherboard on demand. This introduces a

4:41tiny delay, but makes massive models

4:43runnable on normal computers. Updates to

4:46AMD software also delivered a 60%

4:48performance boost on Radiant cards. This

4:51makes non- Nvidia graphics cards viable

4:53options for local workflows, but

4:55shuttling data across a motherboard bus

4:57still creates a physical bottleneck.

5:00The second major technology changing

5:02this market is unified memory

5:04architecture. In traditional computers,

5:06the central processor and graphics card

5:08maintain separate memory pools. Data

5:11must constantly travel between them over

5:13a constrained connection. Unified memory

5:16eliminates that barrier by placing both

5:18processors on a single memory pool.

5:20Apple proved this concept at the high

5:22end with their custom silicon. A Mac

5:24Studio supports up to 512 GB of memory.

5:28It provides incredible bandwidth

5:29exceeding 1 TBTE per second. That

5:32machine runs giant models completely

5:34inside local memory without quality

5:36loss. But spending $5,000 on Apple

5:38hardware is impossible for most

5:40builders. The real consumer breakthrough

5:42comes from AMD with their upcoming

5:44stricks Halo platform. This hybrid

5:47processor packs 16 high performance

5:49cores alongside a powerful graphics

5:51unit. The chip interfaces directly with

5:53up to 128 GB of fast memory. Because the

5:57memory is unified, you do not need

5:58dedicated video memory. The processor

6:01communicates at 256 GB per second across

6:04a wide interface. [music] This unified

6:06layout completely bypasses the

6:08traditional bottlenecks that plague

6:09regular desktop computers. You can go

6:12into the system settings and dedicate 96

6:14GB entirely to graphics. This allows a

6:17tiny energyefficient computer to run a

6:1970 billion parameter model. Perfectly.

6:22Best of all, the entire platform

6:23operates at roughly 120 W. You no longer

6:26need to purchase noisy industrial power

6:28units or multipleuse graphics cards, but

6:31an extraordinary mathematical

6:32breakthrough attacks the problem from

6:34the algorithmic side.

6:36The third and most radical development

6:38is the arrival of one bit language

6:40models. Researchers recently created a

6:43design known as bitnet B158.

6:46Traditional artificial intelligence

6:47models store weights as complex

6:49floatingoint numbers. Calculating those

6:52values requires massive matrix

6:54multiplication that only specialized

6:56processors can handle. Bitnet changes

6:58this fundamental math by training models

7:00with simple turnary values. Every single

7:03weight in the network is restricted to

7:05negative 1 0 or positive 1. The system

7:08entirely eliminates expensive floating

7:10point multiplication. When a weight is

7:12positive 1, the processor simply adds

7:14the incoming value. When a weight is

7:16negative one, the processor subtracts

7:18it. When a weight is zero, the

7:20calculation is skipped completely. This

7:22simple arithmetic cuts the energy cost

7:25of matrix operations by more than 70

7:27times. Memory consumption drops by up to

7:3090%. A 2 billion parameter model that

7:33normally takes 6 GB shrinks down to 400

7:36MGB. Testers successfully ran this model

7:39on a cheap single board computer at 40

7:42tokens per second. Recent research

7:44confirms this architecture scales

7:46successfully up to massive model sizes.

7:49If models no longer require floatingoint

7:51multiplication, expensive graphics cards

7:53lose their primary purpose. By 2028, the

7:56semiconductor memory shortage will

7:58finally end. That price reduction will

8:00coincide with production ready 1-bit

8:02models running on standard processors.

8:05You will no longer need power- hungry

8:07graphics cards to experience top tier

8:09intelligence locally. Running a massive

8:11model on your personal computer will

8:13soon become incredibly simple. Which

8:15large language model are you most

8:17excited to run locally once hardware

8:19barriers disappear?

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.