Full transcript
Intro
0:01Hey everybody, it's Lon Siden. Over the
0:03last couple of weeks, I've been delving
0:04into local AI. It's a new interest area
0:07for me and for many of you given my
0:09analytics on this topic, and I wanted to
0:11find the least expensive 32 gigabyte GPU
0:14out there that you can get as a
0:16consumer. Now, we did look at an Intel
0:18GPU with those specs a few weeks ago,
0:20but that was about 1,300 bucks. For
0:23about half that price, you can pick up a
0:25data center pole like this one. And this
0:28is a Tesla V100. It's not made by Tesla,
0:31the car company, but rather Nvidia. And
0:33although this does not have the bells
0:35and whistles that modern RTX cards have
0:38for video generation and images and
0:40stuff, it is actually very well suited
0:42for local language models at the moment.
0:44This might be a great way to run a
0:46fairly robust local model on consumer
0:49hardware, but there are some steps
0:50involved in getting this thing to work.
0:53And in today's video, we're going to get
0:54all of those steps completed, hook it up
0:56to a PC and get the language model
0:59working. Now, I do want to let you know
1:00in the interest of full disclosure that
1:02I paid for the GPU here with my own
1:04funds along with all of the parts that
1:05are going to be attached to it. There is
1:07a mini PC that will appear in a little
1:09bit that came in free of charge from GMK
1:11Tech. However, no other compensation was
1:13received. All the opinions you're about
1:15to hear are my own, and no one has
1:17reviewed or approved what you're about
1:18to see before it was uploaded. Now, one
Getting a Tesla V100 to work on a desktop
1:20of the challenges with these Tesla cards
1:22is that there is no active cooling on
1:24them because in a data center, the case
1:27that it's installed in will blow air
1:29through the card to keep it cool. So,
1:31although there's no fan, this is not a
1:33passive cooling solution. So, what I
1:35picked up on AliExpress here was a fan,
1:37a blower that we can put on the end of
1:39it. I'm not sure how loud this is going
1:41to be, but we will find out. And so,
1:43what we're going to do is screw this on
1:45to the back of the GPU here. And when
1:48it's attached completely, it's going to
1:50look like so. And this fan should be
1:52enough to keep the card cool even under
1:55load. Now, one of the limitations I have
1:58right now is that I don't have a desktop
2:01computer that can run all of this stuff.
2:02So, we're going to run this as an
2:04external GPU, and we're going to connect
Oculink Dock and Power Connections
2:06it up to this mini PC here using its
2:09Oculink connection, which basically
2:11gives you an external PCI Express slot.
2:14And that slot lives on top of this
2:16Oculink external enclosure. And this is
2:20from AAR. I guess that's how you
2:22pronounce their title there. It's not
2:24very expensive. It's an 800 watt power
2:26supply with a PCI Express slot on the
2:28top. And it has three power outputs,
2:31which is going to be important because
2:34we have to power not only the GPU here
2:36with two of these outputs, but the fan
2:39has to get powered by the third one. And
2:42so to do all of that, I had to buy a
2:44whole bunch of cables here. So the first
2:45cable I had to get was this power supply
2:48cable uh from AliExpress. You can find
2:50these on Amazon also. And this will
2:52adapt the power port on the back of the
2:56Tesla card here, which is a kind of data
2:58center port to a consumer grade power
3:00port here. So we're going to use this.
3:03And then the external enclosure came
3:06with three power cables. So we will be
3:08connecting those power cables up to each
3:10of these. and plugging them in. And then
3:13the fan I had to buy a whole bunch of
3:14extra cables for. So I got this uh Molex
3:18to uh fan connector here. And then I
3:21realized that the two ends were both uh
3:23female or male. So I had to get a gender
3:25changer here. And then we'll plug that
3:28into the third power adapter there. And
3:30hopefully that will keep our fan going.
3:32And the fan will not have a variable
3:35rate on it. So it's going to be blowing
3:36at full blast all the time. And I've
3:39been following some of the folks on
3:40YouTube who have all these crazy ways to
3:42cool this card off. So, there are ways
3:43in which you can put a variable rate fan
3:45on here and stuff, but today we're
3:47shooting for the minimally viable
3:49product, which is getting this thing to
3:50work, and then over time, we'll find
3:52ways to maybe make that fan a little
3:54more efficient. So, this is definitely
3:56not a plug-and-play kind of solution
3:58here. There's a lot to it, but I think
3:59we've got enough here to make it work,
4:01hopefully. So, we're going to get
4:03started on that. So, what I'm going to
4:04do first is get the fan attached and
4:06then we'll get it hooked up to the GPU
4:09enclosure here and make sure that the
4:11fan will operate and then we'll get the
4:12computer out and see if we can get this
Powering the cooling system
4:14thing to work. All right, let's see if
4:15we can get the fan to work first here. I
4:17do have the enclosure here powered up.
4:19Believe it or not, this thing idles with
4:21nothing in it at 9 watts. It might be
4:22the fan inside perhaps. I don't know.
4:24Um, so we're going to connect up the
4:26crazy fan cable contraption that I have
4:28here. And this video is either going to
4:30be really quick or really long depending
4:32on how the outcome of this part goes.
4:34But there we go. The fan has been
4:36connected and
4:38it is blowing. And it's not all that
4:40loud. And it looks like most of the air
4:43is getting out of here. So hopefully
4:44it's enough. One of the things you'll
4:46run into, of course, is this card will
4:47throttle if that fan is not able to keep
4:50it cool. And it's going to run only at
4:53one speed now, which is the fastest
4:55speed it's at at the moment. So we'll
4:56see. But at least right now, the fan is
4:58working off the power supply, and I'm
5:00able now to connect up the rest of the
5:02rig here. As the fan is going, we're now
5:05consuming about 12 watts. And that's
5:06because this fan is rated for about 3
5:09and a/4 watts of power consumption. So,
5:12the next step is to get our card mounted
5:14up and connected to the computer. All
eGPU Installation
5:16right, so why don't we get the card
5:17hooked up here? So, we will go ahead and
5:19just insert it into the PCI Express slot
5:22and snap it in. And then I will close
5:25the little screw here to get it locked
5:27down and hopefully stable, which it
5:29looks like it is. So that is good. That
5:32is all attached up. And then the next
5:34step, of course, is getting the power
5:35supply attached here. So I'll spin it
5:37around to the back so you can see how
5:38this goes. I was concerned that the
5:41power cord would get blocked by the fan,
5:42but they cut a little notch into the 3D
5:46printed shroud here. So hopefully we can
5:48keep these wires from getting in mixed
5:51up with the fan as we're running here.
5:53But I'm going to connect up the first
5:54cable here and attach that into one of
5:57our power sources here. And then we'll
5:59do the other one. And then once all of
6:02this is together, we can power the eGPU
6:04back up. This does have its own power
6:06switch, which can be helpful. And we'll
6:08go ahead and get this one going. And
6:10then we can get the computer into the
6:12mix and hopefully install some drivers
6:15and get this thing up and running. So,
6:17let's get that attached. And now we've
6:19got this thing fully armed and
6:20operational. And I guess I could power
6:23it up real quick and make sure that it
6:24doesn't explode. So, why don't we try
6:26that and just make sure all is good. The
6:28fan is blowing and we're consuming about
6:3130 watts now. So, it looks like the GPU
6:33is drawing some power. But, of course,
6:35there's no computer hooked up to it just
6:36yet. So, what the next step will be is
6:39to get that computer going and then
6:40we'll hunt around for some drivers and
6:42then see what we can do. All right, we
Linux vs. Windows
6:44got everything up and running now.
6:45However, I was not able to get this to
6:47work under Windows. I was able to
6:49install the drivers for this card. the
6:51Nvidia data center drivers for Windows
6:5311. The card was detected, but I kept
6:55getting these code 10 errors and I think
6:57it might be due to the nature of this
6:59Oculink connection and maybe some of the
7:01BIOS interactions. So, what I ended up
7:03doing was heading over, of course, to
7:05Linux. And what I've been doing on Linux
ChatGPT Assistance with Configuration
7:08with these GPU experiments is having the
7:11codeex app from Chat GPT guide me
7:14through the setup process. And so far,
7:17everything looks good. What it's doing
7:19right now is just finishing up getting a
7:21couple of models installed that we can
7:22use to test and then we'll be able to
7:24see what the performance of this setup
7:26is looking like. And I really like using
7:29uh chat GPT codecs as a companion or
7:32assistant now in getting some hardware
7:34configurations up and running quickly
7:36under Linux because it is able to get
7:39the latest drivers, get through all of
7:41the crazy little configuration gotchas
7:43that you run into, and it is very easy
7:46to get started. And what's nice about
7:47this is that when it's done, you can ask
7:49it to issue you a report to tell you
7:52exactly all the different things that it
7:53did. So you have a reference for future
7:56configuration changes or if you wanted
7:57to install it on another machine. So let
7:59me let this finish up and when it's
8:01done, we'll see what kind of performance
Qwen 3.8 27B Test
8:02we can get out of this. All right, we
8:03are up and running and what you're
8:05looking at here is the Llama CPP web
8:07interface and that is what is serving
8:10the models on this computer. Right now I
8:12have Quen 3.8 installed and it is
8:15loaded. This is the 27 billion parameter
8:18model. This is a dense model which means
8:20it should be the slower of the two that
8:22we're going to test here. I've set the
8:24reasoning to medium as this model tends
8:26to do a lot of reasoning. And what I'm
8:28going to do is give it a PDF document
8:30here and ask it to summarize this. And
8:34I'm going to also ask it um also tell me
8:38what is the ATSC's position on 5G TV.
8:42And that's what this document is about.
8:44So, we'll head go ahead and hit the
8:46enter key here, and it will do its
8:48prefill, which means it's going to load
8:50that document into its context. I have a
8:5396,000 token context on here, and it's
8:55all fitting nicely on the GPU. Right
8:58now, I'm seeing an output speed of about
9:0050 tokens per second. This is certainly
9:02a lot faster than the 32 gigabyte Intel
9:05card we looked at a few weeks ago. So,
9:08it's definitely outputting with this
9:10dense model rather quickly. You can see
9:11how much reasoning goes on even on the
9:14medium setting here. And I found with
9:15these smaller models that the reasoning
9:17does help them get more accurate because
9:20they do spend a little more time
9:21checking themselves before they start
9:23outputting something on screen. So that
9:26is a pretty good output here. And now
9:28we've got the actual thing popping out
9:30and we're running at about 45 tokens per
9:32second. We're drawing about 265 watts
9:35right now out of the uh GPU enclosure
9:38here. And all is good so far. and we'll
9:41let this run out. I will test offline
9:43when we fill up that context length a
9:45little bit more to see if it slows down
9:47at all. So now I've loaded in a much
9:49larger document about 60 or 70 pages
9:52worth very dense text and that uh
9:55document and the reasoning here is now
9:58accounting for about 77%
10:01of the context window. So things will
10:03slow down a bit. I will run my thermal
10:05test here again. It looks like we are
10:06still at full power and we are
10:08generating now with our uh context at
10:11almost 80% about 30 tokens per second.
10:14So not bad and certainly a lot faster
10:17with a smaller job but it is getting the
10:19job done here with a very large document
10:22loaded. And what I had to do is look
10:24through I'll pull it up here. Uh this
10:26FCC uh filing from earlier in the year
10:29or last year uh where Tyler the the
10:32antenna man and I went to the FCC and we
10:35were quoted in this document. I wanted
10:36to see how many times we were noted
10:38here. And as you can see it went through
10:40there and found all of it. And we can
10:43query it a little bit more. Uh tell me
10:45about their DRM questions. and it will
10:47go back out and search through what it
10:49has loaded in now to see uh whether it
10:52can answer that question or not. So
10:53certainly it's going to run a little
10:55slower with a larger document, but still
10:57uh very very usable here and it's going
10:59to process that query. It'll take it
11:02about four more seconds. So maybe we'll
11:04just stretch it out a little bit here
11:05and see what it gives us when it is
11:07done. And there you go. So yeah, so it's
11:09going a little bit slower now. We're
11:10down to about 24 tokens per second. But
11:14for me, this is still certainly usable
11:16out of a very dense model with most of
11:19the context length here filled up. And
11:21we'll do one more check of the
11:23temperature here. And we are still
11:25operating at full power without any
11:27thermal throttling with our blower
11:29attached that you likely hear. Let's
11:31take a look now at a faster model. We'll
11:34look next at Gemma 4's mixture of
Gemma-4 26B-A4B Mixture of Experts
11:36experts. All right, here we go. This is
11:38Gemma 4 26B A4B. We're going to give it
11:41all of those PDFs here and see how it
11:43does with that. Right now we're in the
11:45prefill stage where it's ingesting the
11:48prompt and all of the text from those
11:50PDF documents there. It's running at,00
11:53tokens per second on the prefill. So
11:55these models are certainly faster
11:58because they do not have as much density
12:01to how they approach things. they route
12:02to specific areas of the model that
12:05require less parameters to be doing
12:07analysis. As you'll see here when this
12:09starts cranking out, we're going to be
12:11running here in the 80 to 90 tokens per
12:14second range uh versus what we were
12:16seeing before. So certainly a lot
12:17quicker. This model may not always be as
12:20effective as a denser model, but that's
12:21what's great about local AI. You can
12:23find the model that works well for your
12:25particular situation. Uh but here as we
12:27start banging out the response, we're
12:29running at about 90 tokens per second
12:32and I would expect a similar slowdown as
12:34we fill up that context length further.
12:36Uh but certainly you can get some good
12:38performance out of this card here. Now
Video and Image Diffusion Models!
12:40you can do some image diffusion on this
12:42including video generation. I've got
12:44Comfy UI loaded up here with a pretty
12:47recent model Flux 2 Klein 9B. And
12:50although it's slow, it is able to get
12:52images generated. So, right now I've got
12:54this diner image and we're going to have
12:56it update to put a Siberian husky on top
12:59of the motorcycle here on its next run.
13:01We're about halfway through it. It takes
13:03about a minute for it to generate an
13:05image, but when it's done here, we
13:08should see that husky light up there.
13:10We'll just let it finish up here and
13:11I'll show you something else I made with
13:12it. So, there we go. And boom, it is all
13:16set. So, that's pretty cool. And then it
13:18also did video generation. And I loaded
13:20up the new LTX 2.5 model. And to my
13:24surprise, it did run and it looks great.
13:27It's just not very fast. So, this little
13:29clip here took about 3 and 1/2 minutes
13:32or so to generate 5 seconds of video.
13:34And it gets a lot slower the more that
13:37you add in length to the video here. But
13:40it does look pretty cool. By comparison,
13:43this same clip generated uh in about a
13:46minute and a half on my Intel B70. Intel
13:49recently optimized Comfy UI for uh LTX
13:532.5 and also another model that I like
13:55called MiniMax H3 and I covered that in
13:57a prior video. So, if you had a newer
13:59Nvidia card, this will spit out a lot
14:01quicker, but if you have patience, you
14:04can generate video and images on your
14:07pulled GPU here. Now, there are some
14:09gotchas to using a data center GPU like
14:12the one we're playing around with here.
14:14This V100 is from 2017. It's almost 10
14:17years old, so there's been a lot of
14:20developments on the hardware side, and
14:22many new models will be expecting
14:24hardware features that this GPU simply
14:27lacks. We're okay right now for language
14:29models, like you just saw here. But if
14:31you wanted to get into video and image
14:33generation, you will find some stuff
14:35that works, but a lot of the newer stuff
14:37that everyone's going to be talking
14:38about is probably going to not run on
14:41this hardware given that it lacks some
14:42of the hardware features that those
14:45models expect. So, just be aware of
14:46that. You may not get all that many
14:48years out of this, but for now, if you
14:50were just looking for having some local
14:52AI brain on your network with a 32 gig
14:55card that can fit 96,000 tokens of
14:57context on board with its 32 gigs of
15:00memory right now, one of these pulled
15:02Tesla cards is probably the best deal
15:04you're going to get. That'll do it for
15:06now. Until next time, this is Lon Cybin.
15:07Thanks for watching.