Free YouTube Transcribe

Video transcript

The Cheapest 32GB Nvidia GPU You Can Buy for Local AI (Tesla V100)

Lon.TV · 3,147 words · 15 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

Intro

0:01Hey everybody, it's Lon Siden. Over the

0:03last couple of weeks, I've been delving

0:04into local AI. It's a new interest area

0:07for me and for many of you given my

0:09analytics on this topic, and I wanted to

0:11find the least expensive 32 gigabyte GPU

0:14out there that you can get as a

0:16consumer. Now, we did look at an Intel

0:18GPU with those specs a few weeks ago,

0:20but that was about 1,300 bucks. For

0:23about half that price, you can pick up a

0:25data center pole like this one. And this

0:28is a Tesla V100. It's not made by Tesla,

0:31the car company, but rather Nvidia. And

0:33although this does not have the bells

0:35and whistles that modern RTX cards have

0:38for video generation and images and

0:40stuff, it is actually very well suited

0:42for local language models at the moment.

0:44This might be a great way to run a

0:46fairly robust local model on consumer

0:49hardware, but there are some steps

0:50involved in getting this thing to work.

0:53And in today's video, we're going to get

0:54all of those steps completed, hook it up

0:56to a PC and get the language model

0:59working. Now, I do want to let you know

1:00in the interest of full disclosure that

1:02I paid for the GPU here with my own

1:04funds along with all of the parts that

1:05are going to be attached to it. There is

1:07a mini PC that will appear in a little

1:09bit that came in free of charge from GMK

1:11Tech. However, no other compensation was

1:13received. All the opinions you're about

1:15to hear are my own, and no one has

1:17reviewed or approved what you're about

1:18to see before it was uploaded. Now, one

Getting a Tesla V100 to work on a desktop

1:20of the challenges with these Tesla cards

1:22is that there is no active cooling on

1:24them because in a data center, the case

1:27that it's installed in will blow air

1:29through the card to keep it cool. So,

1:31although there's no fan, this is not a

1:33passive cooling solution. So, what I

1:35picked up on AliExpress here was a fan,

1:37a blower that we can put on the end of

1:39it. I'm not sure how loud this is going

1:41to be, but we will find out. And so,

1:43what we're going to do is screw this on

1:45to the back of the GPU here. And when

1:48it's attached completely, it's going to

1:50look like so. And this fan should be

1:52enough to keep the card cool even under

1:55load. Now, one of the limitations I have

1:58right now is that I don't have a desktop

2:01computer that can run all of this stuff.

2:02So, we're going to run this as an

2:04external GPU, and we're going to connect

Oculink Dock and Power Connections

2:06it up to this mini PC here using its

2:09Oculink connection, which basically

2:11gives you an external PCI Express slot.

2:14And that slot lives on top of this

2:16Oculink external enclosure. And this is

2:20from AAR. I guess that's how you

2:22pronounce their title there. It's not

2:24very expensive. It's an 800 watt power

2:26supply with a PCI Express slot on the

2:28top. And it has three power outputs,

2:31which is going to be important because

2:34we have to power not only the GPU here

2:36with two of these outputs, but the fan

2:39has to get powered by the third one. And

2:42so to do all of that, I had to buy a

2:44whole bunch of cables here. So the first

2:45cable I had to get was this power supply

2:48cable uh from AliExpress. You can find

2:50these on Amazon also. And this will

2:52adapt the power port on the back of the

2:56Tesla card here, which is a kind of data

2:58center port to a consumer grade power

3:00port here. So we're going to use this.

3:03And then the external enclosure came

3:06with three power cables. So we will be

3:08connecting those power cables up to each

3:10of these. and plugging them in. And then

3:13the fan I had to buy a whole bunch of

3:14extra cables for. So I got this uh Molex

3:18to uh fan connector here. And then I

3:21realized that the two ends were both uh

3:23female or male. So I had to get a gender

3:25changer here. And then we'll plug that

3:28into the third power adapter there. And

3:30hopefully that will keep our fan going.

3:32And the fan will not have a variable

3:35rate on it. So it's going to be blowing

3:36at full blast all the time. And I've

3:39been following some of the folks on

3:40YouTube who have all these crazy ways to

3:42cool this card off. So, there are ways

3:43in which you can put a variable rate fan

3:45on here and stuff, but today we're

3:47shooting for the minimally viable

3:49product, which is getting this thing to

3:50work, and then over time, we'll find

3:52ways to maybe make that fan a little

3:54more efficient. So, this is definitely

3:56not a plug-and-play kind of solution

3:58here. There's a lot to it, but I think

3:59we've got enough here to make it work,

4:01hopefully. So, we're going to get

4:03started on that. So, what I'm going to

4:04do first is get the fan attached and

4:06then we'll get it hooked up to the GPU

4:09enclosure here and make sure that the

4:11fan will operate and then we'll get the

4:12computer out and see if we can get this

Powering the cooling system

4:14thing to work. All right, let's see if

4:15we can get the fan to work first here. I

4:17do have the enclosure here powered up.

4:19Believe it or not, this thing idles with

4:21nothing in it at 9 watts. It might be

4:22the fan inside perhaps. I don't know.

4:24Um, so we're going to connect up the

4:26crazy fan cable contraption that I have

4:28here. And this video is either going to

4:30be really quick or really long depending

4:32on how the outcome of this part goes.

4:34But there we go. The fan has been

4:36connected and

4:38it is blowing. And it's not all that

4:40loud. And it looks like most of the air

4:43is getting out of here. So hopefully

4:44it's enough. One of the things you'll

4:46run into, of course, is this card will

4:47throttle if that fan is not able to keep

4:50it cool. And it's going to run only at

4:53one speed now, which is the fastest

4:55speed it's at at the moment. So we'll

4:56see. But at least right now, the fan is

4:58working off the power supply, and I'm

5:00able now to connect up the rest of the

5:02rig here. As the fan is going, we're now

5:05consuming about 12 watts. And that's

5:06because this fan is rated for about 3

5:09and a/4 watts of power consumption. So,

5:12the next step is to get our card mounted

5:14up and connected to the computer. All

eGPU Installation

5:16right, so why don't we get the card

5:17hooked up here? So, we will go ahead and

5:19just insert it into the PCI Express slot

5:22and snap it in. And then I will close

5:25the little screw here to get it locked

5:27down and hopefully stable, which it

5:29looks like it is. So that is good. That

5:32is all attached up. And then the next

5:34step, of course, is getting the power

5:35supply attached here. So I'll spin it

5:37around to the back so you can see how

5:38this goes. I was concerned that the

5:41power cord would get blocked by the fan,

5:42but they cut a little notch into the 3D

5:46printed shroud here. So hopefully we can

5:48keep these wires from getting in mixed

5:51up with the fan as we're running here.

5:53But I'm going to connect up the first

5:54cable here and attach that into one of

5:57our power sources here. And then we'll

5:59do the other one. And then once all of

6:02this is together, we can power the eGPU

6:04back up. This does have its own power

6:06switch, which can be helpful. And we'll

6:08go ahead and get this one going. And

6:10then we can get the computer into the

6:12mix and hopefully install some drivers

6:15and get this thing up and running. So,

6:17let's get that attached. And now we've

6:19got this thing fully armed and

6:20operational. And I guess I could power

6:23it up real quick and make sure that it

6:24doesn't explode. So, why don't we try

6:26that and just make sure all is good. The

6:28fan is blowing and we're consuming about

6:3130 watts now. So, it looks like the GPU

6:33is drawing some power. But, of course,

6:35there's no computer hooked up to it just

6:36yet. So, what the next step will be is

6:39to get that computer going and then

6:40we'll hunt around for some drivers and

6:42then see what we can do. All right, we

Linux vs. Windows

6:44got everything up and running now.

6:45However, I was not able to get this to

6:47work under Windows. I was able to

6:49install the drivers for this card. the

6:51Nvidia data center drivers for Windows

6:5311. The card was detected, but I kept

6:55getting these code 10 errors and I think

6:57it might be due to the nature of this

6:59Oculink connection and maybe some of the

7:01BIOS interactions. So, what I ended up

7:03doing was heading over, of course, to

7:05Linux. And what I've been doing on Linux

ChatGPT Assistance with Configuration

7:08with these GPU experiments is having the

7:11codeex app from Chat GPT guide me

7:14through the setup process. And so far,

7:17everything looks good. What it's doing

7:19right now is just finishing up getting a

7:21couple of models installed that we can

7:22use to test and then we'll be able to

7:24see what the performance of this setup

7:26is looking like. And I really like using

7:29uh chat GPT codecs as a companion or

7:32assistant now in getting some hardware

7:34configurations up and running quickly

7:36under Linux because it is able to get

7:39the latest drivers, get through all of

7:41the crazy little configuration gotchas

7:43that you run into, and it is very easy

7:46to get started. And what's nice about

7:47this is that when it's done, you can ask

7:49it to issue you a report to tell you

7:52exactly all the different things that it

7:53did. So you have a reference for future

7:56configuration changes or if you wanted

7:57to install it on another machine. So let

7:59me let this finish up and when it's

8:01done, we'll see what kind of performance

Qwen 3.8 27B Test

8:02we can get out of this. All right, we

8:03are up and running and what you're

8:05looking at here is the Llama CPP web

8:07interface and that is what is serving

8:10the models on this computer. Right now I

8:12have Quen 3.8 installed and it is

8:15loaded. This is the 27 billion parameter

8:18model. This is a dense model which means

8:20it should be the slower of the two that

8:22we're going to test here. I've set the

8:24reasoning to medium as this model tends

8:26to do a lot of reasoning. And what I'm

8:28going to do is give it a PDF document

8:30here and ask it to summarize this. And

8:34I'm going to also ask it um also tell me

8:38what is the ATSC's position on 5G TV.

8:42And that's what this document is about.

8:44So, we'll head go ahead and hit the

8:46enter key here, and it will do its

8:48prefill, which means it's going to load

8:50that document into its context. I have a

8:5396,000 token context on here, and it's

8:55all fitting nicely on the GPU. Right

8:58now, I'm seeing an output speed of about

9:0050 tokens per second. This is certainly

9:02a lot faster than the 32 gigabyte Intel

9:05card we looked at a few weeks ago. So,

9:08it's definitely outputting with this

9:10dense model rather quickly. You can see

9:11how much reasoning goes on even on the

9:14medium setting here. And I found with

9:15these smaller models that the reasoning

9:17does help them get more accurate because

9:20they do spend a little more time

9:21checking themselves before they start

9:23outputting something on screen. So that

9:26is a pretty good output here. And now

9:28we've got the actual thing popping out

9:30and we're running at about 45 tokens per

9:32second. We're drawing about 265 watts

9:35right now out of the uh GPU enclosure

9:38here. And all is good so far. and we'll

9:41let this run out. I will test offline

9:43when we fill up that context length a

9:45little bit more to see if it slows down

9:47at all. So now I've loaded in a much

9:49larger document about 60 or 70 pages

9:52worth very dense text and that uh

9:55document and the reasoning here is now

9:58accounting for about 77%

10:01of the context window. So things will

10:03slow down a bit. I will run my thermal

10:05test here again. It looks like we are

10:06still at full power and we are

10:08generating now with our uh context at

10:11almost 80% about 30 tokens per second.

10:14So not bad and certainly a lot faster

10:17with a smaller job but it is getting the

10:19job done here with a very large document

10:22loaded. And what I had to do is look

10:24through I'll pull it up here. Uh this

10:26FCC uh filing from earlier in the year

10:29or last year uh where Tyler the the

10:32antenna man and I went to the FCC and we

10:35were quoted in this document. I wanted

10:36to see how many times we were noted

10:38here. And as you can see it went through

10:40there and found all of it. And we can

10:43query it a little bit more. Uh tell me

10:45about their DRM questions. and it will

10:47go back out and search through what it

10:49has loaded in now to see uh whether it

10:52can answer that question or not. So

10:53certainly it's going to run a little

10:55slower with a larger document, but still

10:57uh very very usable here and it's going

10:59to process that query. It'll take it

11:02about four more seconds. So maybe we'll

11:04just stretch it out a little bit here

11:05and see what it gives us when it is

11:07done. And there you go. So yeah, so it's

11:09going a little bit slower now. We're

11:10down to about 24 tokens per second. But

11:14for me, this is still certainly usable

11:16out of a very dense model with most of

11:19the context length here filled up. And

11:21we'll do one more check of the

11:23temperature here. And we are still

11:25operating at full power without any

11:27thermal throttling with our blower

11:29attached that you likely hear. Let's

11:31take a look now at a faster model. We'll

11:34look next at Gemma 4's mixture of

Gemma-4 26B-A4B Mixture of Experts

11:36experts. All right, here we go. This is

11:38Gemma 4 26B A4B. We're going to give it

11:41all of those PDFs here and see how it

11:43does with that. Right now we're in the

11:45prefill stage where it's ingesting the

11:48prompt and all of the text from those

11:50PDF documents there. It's running at,00

11:53tokens per second on the prefill. So

11:55these models are certainly faster

11:58because they do not have as much density

12:01to how they approach things. they route

12:02to specific areas of the model that

12:05require less parameters to be doing

12:07analysis. As you'll see here when this

12:09starts cranking out, we're going to be

12:11running here in the 80 to 90 tokens per

12:14second range uh versus what we were

12:16seeing before. So certainly a lot

12:17quicker. This model may not always be as

12:20effective as a denser model, but that's

12:21what's great about local AI. You can

12:23find the model that works well for your

12:25particular situation. Uh but here as we

12:27start banging out the response, we're

12:29running at about 90 tokens per second

12:32and I would expect a similar slowdown as

12:34we fill up that context length further.

12:36Uh but certainly you can get some good

12:38performance out of this card here. Now

Video and Image Diffusion Models!

12:40you can do some image diffusion on this

12:42including video generation. I've got

12:44Comfy UI loaded up here with a pretty

12:47recent model Flux 2 Klein 9B. And

12:50although it's slow, it is able to get

12:52images generated. So, right now I've got

12:54this diner image and we're going to have

12:56it update to put a Siberian husky on top

12:59of the motorcycle here on its next run.

13:01We're about halfway through it. It takes

13:03about a minute for it to generate an

13:05image, but when it's done here, we

13:08should see that husky light up there.

13:10We'll just let it finish up here and

13:11I'll show you something else I made with

13:12it. So, there we go. And boom, it is all

13:16set. So, that's pretty cool. And then it

13:18also did video generation. And I loaded

13:20up the new LTX 2.5 model. And to my

13:24surprise, it did run and it looks great.

13:27It's just not very fast. So, this little

13:29clip here took about 3 and 1/2 minutes

13:32or so to generate 5 seconds of video.

13:34And it gets a lot slower the more that

13:37you add in length to the video here. But

13:40it does look pretty cool. By comparison,

13:43this same clip generated uh in about a

13:46minute and a half on my Intel B70. Intel

13:49recently optimized Comfy UI for uh LTX

13:532.5 and also another model that I like

13:55called MiniMax H3 and I covered that in

13:57a prior video. So, if you had a newer

13:59Nvidia card, this will spit out a lot

14:01quicker, but if you have patience, you

14:04can generate video and images on your

14:07pulled GPU here. Now, there are some

14:09gotchas to using a data center GPU like

14:12the one we're playing around with here.

14:14This V100 is from 2017. It's almost 10

14:17years old, so there's been a lot of

14:20developments on the hardware side, and

14:22many new models will be expecting

14:24hardware features that this GPU simply

14:27lacks. We're okay right now for language

14:29models, like you just saw here. But if

14:31you wanted to get into video and image

14:33generation, you will find some stuff

14:35that works, but a lot of the newer stuff

14:37that everyone's going to be talking

14:38about is probably going to not run on

14:41this hardware given that it lacks some

14:42of the hardware features that those

14:45models expect. So, just be aware of

14:46that. You may not get all that many

14:48years out of this, but for now, if you

14:50were just looking for having some local

14:52AI brain on your network with a 32 gig

14:55card that can fit 96,000 tokens of

14:57context on board with its 32 gigs of

15:00memory right now, one of these pulled

15:02Tesla cards is probably the best deal

15:04you're going to get. That'll do it for

15:06now. Until next time, this is Lon Cybin.

15:07Thanks for watching.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.