Free YouTube Transcribe

Video transcript

Most Optimization Advice Misses the REAL Problem

SimonDev · 3,882 words · 18 min read

Want to search this transcript, jump the video from any line, or download it as TXT, SRT, or VTT?

Open in the transcript tool

Full transcript

0:00My first major project was a game called

0:02Prototype. And before it was green lit,

0:04we had to produce a vertical slice of

0:06gameplay in about three months.

0:08Performance quickly became the biggest

0:10issue. Nobody had any idea how to get it

0:12from a couple FPS up to a showable 30.

0:15And so I got a trial by fire

0:17introduction to optimization and

0:20continued in that role afterwards and on

0:22other projects later in my career.

0:24Optimization is an incredibly deep

0:26subject, as I quickly found out, and

0:29there are so many ways to approach it. I

0:31don't claim to be an expert, but I've

0:33been lucky enough to have gotten

0:35opportunities to play a lot with the

0:37subject. Let's start with a common and

0:39vague problem intentionally. Let's draw

0:42a whole whack of trees in the scene. So,

0:44here we go. They're coming in. And as

0:46they come in, oh, look, the frame rate

0:49is dropping. Let's pretend you know

0:51nothing like I did. How do you make this

0:53faster? Because you know very well that

0:56games do render jillions of trees, among

0:59other things, quite successfully. So,

1:01it's absolutely possible we're just

1:03doing something seriously wrong here.

1:05Often, it starts with a couple simple

1:07questions. How fast does this need to

1:09be? And on what? What's your target

1:11frame rate? You going for 30 fps on a

1:14low-end phone, 120 fps on a PS5? In

1:18essence, what's your budget? This is 1

1:20second. And so if you're running at 60

1:23fps, then it means that this little bar

1:25here, 16 and 2/3 milliseconds, this is

1:28what you get to render your scene. At a

1:30studio, you'd have to divide this time

1:32obviously with other things that want

1:34part of that time, too, because I assume

1:36your entire game isn't just about

1:38drawing trees. So, you have to share

1:39your frame budget, both CPU and GPU,

1:42with other stuff. In this little frame

1:44breakdown study on the graphics of Doom

1:46Eternal, you can see that a lot of

1:49things go into drawing a frame. You've

1:50got your GPU skinning to animate

1:52characters. There's also shadow mapping.

1:55They've got a depth prepass high zed

1:57screen space, ambient occlusion,

1:59particles, post effects, UI, whatever.

2:01There's a bunch of stuff in here. You

2:03don't have to understand it all. But

2:04during development, these would

2:06absolutely have budgets on the target

2:08hardware, how much frame time they can

2:10have in order to hit the target frame

2:12rate. Now, we need to start our

2:14investigation. Manually taking timings

2:16can get you surprisingly far. or better

2:19yet, some sort of tool to examine

2:20performance, a profiler. Depending on

2:23what you're profiling, let's

2:24oversimplify and restrict ourselves to

2:26CPU or GPU problems. Then you'll need a

2:29way of getting information about

2:30slowdowns. Some common choices would

2:33include tools by Intel, AMD, or Nvidia,

2:36stuff like Vtune, or Insight. I found a

2:39nice little list here on this Rust

2:41profiling page by Nicholas Nethercoat,

2:43who did a lot of the work with the Rust

2:45compiler performance for Firefox. We'll

2:48do some top- down analysis. If you're

2:50doing profiling, it'll generally look

2:52like this. That is starting from the

2:54top. Here's where we update all the

2:56things. That breaks up into smaller

2:58chunks where we update the foods and

3:00then we update the blah. And we'll try

3:02to see where all the time is being

3:03spent. You can see why this is called

3:05top- down analysis because you can drill

3:07down further and figure out where your

3:09overall time is being spent. bottom up

3:11by contrast. We'll look again at the

3:13frame, but we'll look at functions being

3:15used everywhere, like this matrix

3:17update. And by making those fast, you

3:19make everything fast. This is great for

3:21finding those functions where you really

3:23want to squeeze more performance out of

3:25them. In our actual case, let's figure

3:27out what's wrong. So, in the top left,

3:29I've got the frame timings running, both

3:31the CPU and GPU time, because this

3:34problem can potentially span both. We

3:36can see the CPU sucks more than the GPU.

3:38So, let's use a profiler to see why. So,

3:41we fire up a profiler of some sort. They

3:43all look different, but present similar

3:44information. Because I'm a lazy man,

3:46we're using JavaScript and Chrome's

3:48DevTool profiler, but it'll present a

3:51breakdown similar to others. So, we can

3:53see from this breakdown, we're spending

3:55well all of the time on the CPU on draw

3:58commands, which isn't exactly a

4:00surprise. We're not even drawing that

4:01many. So, we want to make this faster.

4:04How do we do that? The general steps to

4:06fix any performance problem look

4:08something like this. You establish a

4:10budget that's already done. We're

4:11dedicating our whole frame to trees. So

4:14the next steps would be well, we'll fix

4:16this problem, make it nice and small,

4:18and then if performance still sucks, we

4:20loop around and figure out what the next

4:22biggest issue is. Let's start with a

4:24really simple example. You've got this

4:26function here, really slow, and your

4:28profiler is telling you this thing is a

4:30problem. What's your first move? I mean,

4:32we could take a look at it and maybe

4:34we'll get really lucky and there's a

4:35bunch of really, really stupid things

4:37going on in there. Remove those and

4:39voila, some of our easiest performance

4:41problems are solved. If we take our tree

4:44example, well, maybe someone

4:45intentionally wrote this super poorly

4:47for YouTube video reasons. Now, we can

4:49share materials, share underlying

4:51geometries. Those begin to give you some

4:53modest gains. We've bumped up our number

4:55of trees for the same frame rate. Looks

4:57like we're hitting about 7 to 800 in

4:59this nonsense example. Even better is

5:01using GPU instancing. That's essentially

5:04telling the GPU, we're going to draw the

5:06same tree a whole bunch of times. Go do

5:08it. That gets us up to 30,000ish trees.

5:12We're also spending almost no time on

5:13the CPU now. So, we're probably

5:15bottlenecked on the GPU. Once that

5:17absurdly lowanging fruit is out of the

5:20way, now what? Well, let's back up a bit

5:22and think about things in terms of what

5:24does your system, the one you're

5:25optimizing, what does it look like at

5:28its core? Stuff goes in. These are your

5:30inputs to the system. You process them

5:33in some way and then stuff goes out. The

5:35question is then what parts of this

5:37process do you control? Because changes

5:40to this processing step here, they're at

5:42the mercy of your input. So if you

5:44control the input, it can be much more

5:46lucrative to start there instead. When

5:48you're going to change the input, often

5:50it's productive to think about, well,

5:52what does my hardware even want? What

5:55shape should the input take to maximize

5:57hardware utilization? In general,

6:00hardware likes a few things. This is a

6:02sweeping generalization, but it likes

6:04giant contiguous chunks of memory and it

6:06likes predictability. Let's say we're

6:09talking about the CPU and we're doing

6:10work with things in memory. If you

6:12process things in a contiguous chunk of

6:14memory, you process this chunk here,

6:16then afterwards you process this chunk

6:18here and then this one here, well, the

6:21CPU can boldly predict that you'll need

6:23this chunk here and just have it ready

6:25to roll when you move on to process your

6:27next chunk. Crazy, eh? But if you're

6:29randomly bouncing around memory looking

6:31for bits and pieces, the CPU has no

6:33goddamn idea where you're going to look

6:35next and so eat a big fat cache miss.

6:38Random access memory can be accessed

6:39randomly, but really shouldn't be. Cache

6:42misses are a big deal if you're not sure

6:44what they are. This little box here,

6:46let's say this is the time that it takes

6:48for the CPU to execute a simple

6:50instruction. Now, let's say that you

6:52want to do something that needs to grab

6:54something from memory. Then if it lucks

6:56out and that thing is in the L1 cache,

6:58it takes about this long in relative

7:00terms. But L1 caches are small. So if

7:03it's not there, often the CPU's next

7:06option is to check L2, which takes a bit

7:08longer, but not crazy long. Miss that

7:12though, and now you're reaching out to

7:13L3, and the cost of doing this is

7:15steadily rising. After that, if the data

7:18you're looking for isn't in any of the

7:20caches, or in other words, this is a

7:22cache miss, then you're reaching out to

7:24main memory, which takes roughly forever

7:27in computer terms. This is why cache

7:29misses hurt so much. So, if you lay out

7:31your data in a way that the CPU can

7:33predict what you're going to need next,

7:35you may eat a lot less of these. And

7:38your CPU can even try to predict which

7:40way your code is going to execute. If we

7:42compare these two loops, they're

7:44identical for the most part. So this

7:46branch here on the left predicated on

7:48data in a random order, there's no

7:50stable pattern for the CPU to pick up

7:52on. And so it's going to do so so so

7:55much worse than this one here. Even

7:57though the code itself is the same, the

7:59data is in a predictable pattern. Or

8:02since you control your inputs, you could

8:04simply separate them out into separate

8:06lists. Job done. The differences here

8:08are subtle and easy to miss. If you

8:11imagine your CPU like a little machine

8:13in a factory, you've got your programs

8:15commands coming in as a little assembly

8:17line and the results coming out as a

8:19little assembly line. Inside though, it

8:22doesn't necessarily have to be done in

8:24this order. Imagine the CPU is looking

8:26at these boxes and you can see that this

8:28box here depends on this box here and

8:31this box here, it wants something from

8:33memory. Reading from memory takes

8:35forever in CPU terms. So this whole

8:38factory operation is now blocked. The

8:41CPU wants to keep going. So maybe stuff

8:43gets shuffled a bit. We can execute this

8:45box here while we wait for the other

8:47one. Everything still comes out in the

8:49right order. Let's reset things a bit.

8:52Now imagine this box is a branch. So you

8:55got your if then else. Let's just give

8:57these different colors that you can see

8:58them clearly. Again, the CPU doesn't

9:00want to slow down, right? Figuring out

9:03this if will take a bit of time. So,

9:05while it waits to see which way it

9:07should go, it actually just tries to

9:09guess and keeps going. If they're right,

9:11awesome. If they're wrong, though, they

9:13have to backtrack and start again. It's

9:15a fantastic optimization, but

9:17mispredictions are a problem, among

9:20other things. There's so much to learn

9:22here, it's not realistic to do more than

9:24just put these on your radar. But now

9:26that you're aware, there will be links

9:28in the description, or I'll make more

9:29videos about it later. So, back to that

9:32big question from earlier. What shape

9:34should my data take? So your function

9:36here, it could be a game update loop. It

9:39could be a lot of things here. It

9:41processes a bunch of world objects which

9:43are in some random order. So we iterate,

9:46grab the entity from the world object

9:48and do our little update. From a

9:50performance standpoint, you've got

9:52pointer indirection, you've got virtual

9:54function calls, which means potential

9:55for a lot of cache misses and branch

9:57mispredictions. With some minor changes

9:59to the input, the performance of the

10:01function can substantially increase.

10:04When you're more mindful of data

10:05layouts, like in this example here, we

10:08can burn through updates way faster by

10:10packing them all into contiguous blocks

10:12of memory with no lookups. Now your CPU

10:15has one giant workload to burn through.

10:17The really cool thing here is that since

10:19your data is all tucked away nicely in a

10:21big predictable array, it gets a lot

10:23easier for either you or the compiler to

10:26vectorize this. Many CPUs have the

10:28ability to do SIMD or single instruction

10:31multiple data, meaning you can now

10:33process four or eight or even more

10:35things at once instead of one by one.

10:37The basic idea though is do your stupid

10:40things offline if possible. Once your

10:42data is in order, then you can get your

10:44code in order. Here's a few good places

10:46to start learning about this. And as a

10:48quick aside, if you've ever wondered why

10:50ECS frameworks have emerged so strongly

10:52in recent years, it should make a lot

10:54more sense now. They're in essence a

10:56data first approach. So in our tree

10:59example, we can take our tree and

11:01optimize it. Weld together redundant

11:03vertices. That'll reduce the workload.

11:05We'll reorder the mesh to maximize

11:08vertex cache usage. We'll quantize the

11:10data because you don't need these big

11:12bloated vertices. You can crunch them

11:14down quite a bit. There are a zillion

11:16tricks in this space. Ocahedron mapping

11:18for your normals. An entire tangent

11:20space can be compressed to four bytes.

11:22Positions can be done with half floats

11:24or much better 16- bit integers with

11:27scale and offset. UVs can be similarly

11:29quantized in many cases. And finally, we

11:32can compress the textures using a GPU

11:34compressed format. These might sound

11:36unimportant, but smaller vertices and

11:38textures mean better use of bandwidth

11:40and more crap fits in the cache, meaning

11:43more trees before you hit the GPU's

11:44limits. We look at our forest. It's

11:47getting bigger. You can see in this shot

11:49we've hit over 50,000 trees. So, the

11:51scene is growing quite a bit. One

11:54interesting question arises the first

11:56time you start doing these

11:57optimizations. How fast can this really

11:59go? The first couple of wins are great.

12:02The next few will probably be smaller

12:04and diminishing returns after that.

12:07Sometimes it's helpful to know what the

12:09speed of light is. Or in other words, if

12:11everything was perfect, how fast could

12:13this really go? Because you're not going

12:16to go faster than that. And as you

12:18approach that limit, the amount of time

12:20versus the amount of gain you make will

12:22really start not being worth it. So what

12:25are your other options? Do it smarter.

12:28Or said another way, don't do the things

12:30you were going to do but didn't have to.

12:33That one isn't quite as catchy. There's

12:35really nothing faster than literally

12:37just not doing it. So the first approach

12:39we talked about, we did the same work,

12:41we just made it faster to do that work.

12:44This asks the question, is there a way

12:46to literally produce the same effective

12:48output but do less overall work to get

12:50there? Can we get the same result while

12:53doing less work overall? If you're

12:55walking from A to B, well, you can get

12:57there faster by just doing an allout

12:59sprint. That's doing the same work just

13:01faster. Or you could find shortcuts.

13:04Take a little shortcut through the mall,

13:05fly over the mountain, whatever. Same

13:07destination, but you're shortening the

13:09overall route. This tends to be

13:11extremely domainspecific. So let's look

13:14at our trees. But let's look at the

13:16scene from above. If we highlight this

13:18area here, it's called the viewfrostm.

13:21It corresponds exactly to what you can

13:23see on screen. So these highlighted

13:25trees here, they're visible. While the

13:27rest of these, these are out of view. So

13:30they contribute nothing visually to the

13:31output. But in our naive setup, they

13:34still get sent to the GPU and take up

13:36processing time. So don't draw these.

13:39Simple. Call them. Bam. Instant win.

13:41You've raised the ceiling by quite a

13:42bit. How does the speed of light factor

13:44in here? Previously, it was the maximum

13:47speed that you could do certain

13:48operations like say draw trees. That

13:51hasn't really changed. If I had a

13:53theoretical max of 100 trees before,

13:55that still applies. But the difference

13:57is now they're the ones on screen and

14:00the offscreen ones don't count anymore.

14:02The ceiling still exists, but now you

14:04could have a thousand trees in the world

14:06because only 100 would be on screen at

14:08any given time. There's more avenues to

14:10explore here. In the case of our trees,

14:12we could use some sort of spatial

14:14division to rapidly discard larger

14:16portions of the world before fine grain

14:17checking each tree or take another

14:20approach. Obviously, some geological

14:22features may obstruct our view and hence

14:24anything behind that would be oluded.

14:27Watch my video on occlusion techniques

14:29to understand how much effort goes in

14:31here. But let's say that you somehow get

14:33perfect information. You're able to in

14:35the blink of an eye get a list of every

14:37single visible non-luded tree. Well,

14:40you're still bound to your speed of

14:41light restriction, unfortunately. The

14:44question then becomes, how do you break

14:45the speed of light barrier? Well,

14:47anybody who's been in a grocery store in

14:49the last couple years knows, how do you

14:51make more money without providing a

14:53better product? You don't. You cheat

14:55shamelessly and hope either nobody

14:58notices or cares enough to complain.

15:00Voila, profit. There's some really great

15:02examples over the years of games using

15:04tricks to get away with things. The fog

15:07in Silent Hill obscured your view and

15:09provided a lot of atmosphere and also

15:11ensured devs didn't have to draw too

15:13far. The infuriatingly long elevator

15:16scenes in Mass Effect gave the team time

15:18to banter, let you get to know them

15:20better, and provided cover while they

15:21loaded the scene. And Jack and Daxter

15:24streaming system would apparently have

15:26you trip and fall when you outrun the

15:28game's ability to stream in assets.

15:30That's not from any official docs, but

15:32other games have apparently done the

15:33same thing. Let's turn our attention to

15:36these trees. Here's our tree up close in

15:38all of its glory. And here it is far

15:41away, I guess, not in its glory. Well,

15:43since it's far away, why even bother

15:46drawing a good version of it? We could

15:47substitute a lesser quality version, and

15:50maybe nobody will notice or care. Maybe

15:52I already did. It's so far away. Or

15:54maybe I'm lying. Can you really tell?

15:57This is great for a variety of reasons

15:59which this brilliant YouTuber dives into

16:01in his video on why LODs or level of

16:03detail why those work so well. Beyond

16:06the savings for the smaller simplified

16:08mesh, you reduce quad overdraw. If we

16:11push it really far away, maybe we can

16:12cheat even more. Just putting a picture

16:15of a tree in. I mean, does it really

16:17matter? If this is so far away that it's

16:19being caught in the distance fog, you

16:21can't tell anyway. But the better your

16:23cheating is, the closer you can do it

16:25before anybody cares. Take this simple

16:27cube. This picture kind of sucks, but we

16:30can make it better. Maybe by separating

16:32the diffuse from the normal. Now, these

16:34can be lit nicely with some work. And we

16:36can go even further. Take a look at this

16:38technique from Ryan Brooks called

16:40octahedral impostors. The general idea

16:43is that you take that same tree and take

16:45snapshots from many different angles and

16:47use an octaedral mapping to get you a

16:49nice distribution of angles. Then with

16:51some blending, you can walk around the

16:53object and it kind of reacts like a real

16:55one. Now look at this thing. Even up

16:57close, it actually looks somewhat like

16:59the real thing, but it's all smoke and

17:02mirrors. You can scatter these out in

17:04the field now. They react to lighting.

17:06And at that point, you could even start

17:08playing with post effects like ambient

17:09occlusion or screen space shadows to

17:11make it look even better. One thing I

17:13should call out, be careful of both your

17:15average case and your worst case. If

17:18your average case is great, but every

17:20once in a while you turn the camera and

17:22hit one FPS. That's awful. As Jim kind

17:25of screamed this on Twitter, design your

17:28software to minimize fluctuations in

17:30performance. I'll give you an example.

17:33In a game, let's say you're wandering

17:34around the city. Stuff is happening blah

17:36blah, and then suddenly bombs go off,

17:38explosions, chaos. Let's focus on the

17:41particles in this case. Particles are

17:43expensive. They cause huge amounts of

17:45overdraw. So, multiple explosions go

17:47off. You flood the screen with particles

17:49and frame rate tanks. Except you can't

17:51tank the frame rate. That's not cool.

17:53You have to take this in stride. Well,

17:55you know roughly how fast you can draw

17:57particles, right? You've done your

17:58tests. Speed of light, yada yada. On

18:01prototype back in 2005, 2006, I did a

18:04dumb but surprisingly effective Hail

18:06Mary change. Moments before a big

18:09publisher demo, I put a hacky throttle

18:11on the effects. When all those

18:13explosions went off, we didn't just

18:15spawn everything. We pulled from a

18:17shared budget and damped how many

18:19particles each effect could actually

18:20use. Scaling the density of effects.

18:23This evolved into a game of how do we

18:25distribute the particle budget? Closer,

18:28more important stuff, more budget.

18:30Further away things, less one explosion

18:33or 100, things should stay even. Newer

18:35techniques are much smarter than that,

18:37doing things like dynamically adjusting

18:39the resolution. But the idea is just

18:41don't let the worst case be much

18:43different than your average case. So

18:45overall, that's kind of how optimization

18:48works. We started with our simple scene

18:50unoptimized, and we worked through the

18:52various stages. We got the really

18:54lowhanging fruit out of the way,

18:56improved our input, and made sure to use

18:58the hardware's preferred way of drawing

19:00lots of stuff and saw the scene's

19:02performance grow. After that, we got a

19:04bit smarter about distributing the work,

19:06calling out what was invisible, dividing

19:08the world spatially into regions to

19:10rapidly test. And again, we saw big

19:12improvements. Eventually, you hit

19:14diminishing returns because your

19:16hardware has real verifiable limits. And

19:18that's when you start trying to get

19:20creative. What will people notice? What

19:22can I get away with doing? Do the same

19:24work but better. Do less work. Be

19:26smarter. Cheat and lie. In reality,

19:29optimization work is a mixed bag of

19:30these approaches. But let's not nitpick.

19:33It's generally always that same

19:35approach. You figure out how much you

19:37want to spend on this. You profile and

19:39figure out where your problems are. Find

19:40a problematic bit of code. Make changes

19:43there and profile again. The details

19:45will change as you get more experience.

19:47But let's not over complicate this. If

19:49you came into this video thinking I'd

19:50make you some sort of optimization

19:52monster in 20 minutes or however long

19:54this took, sorry. I've seen even

19:57experienced game devs, guys with a

19:59decade or more experience, blindly

20:02optimizing with no real plan. So, just

20:04getting a good framework as a start. I

20:06don't pretend to be the best or to know

20:08everything. For any devs who have some

20:10interesting war stories of major

20:12performance improvements, comment

20:13sections below. Make sure to share.

20:16Cheers.

Recently added transcripts

Browse the whole transcript library

This transcript was generated from the captions YouTube publishes for this video. Get the transcript of any YouTube video atfreeyoutubetranscribe.com, free, unlimited, no sign-up.