Full transcript
0:00My first major project was a game called
0:02Prototype. And before it was green lit,
0:04we had to produce a vertical slice of
0:06gameplay in about three months.
0:08Performance quickly became the biggest
0:10issue. Nobody had any idea how to get it
0:12from a couple FPS up to a showable 30.
0:15And so I got a trial by fire
0:17introduction to optimization and
0:20continued in that role afterwards and on
0:22other projects later in my career.
0:24Optimization is an incredibly deep
0:26subject, as I quickly found out, and
0:29there are so many ways to approach it. I
0:31don't claim to be an expert, but I've
0:33been lucky enough to have gotten
0:35opportunities to play a lot with the
0:37subject. Let's start with a common and
0:39vague problem intentionally. Let's draw
0:42a whole whack of trees in the scene. So,
0:44here we go. They're coming in. And as
0:46they come in, oh, look, the frame rate
0:49is dropping. Let's pretend you know
0:51nothing like I did. How do you make this
0:53faster? Because you know very well that
0:56games do render jillions of trees, among
0:59other things, quite successfully. So,
1:01it's absolutely possible we're just
1:03doing something seriously wrong here.
1:05Often, it starts with a couple simple
1:07questions. How fast does this need to
1:09be? And on what? What's your target
1:11frame rate? You going for 30 fps on a
1:14low-end phone, 120 fps on a PS5? In
1:18essence, what's your budget? This is 1
1:20second. And so if you're running at 60
1:23fps, then it means that this little bar
1:25here, 16 and 2/3 milliseconds, this is
1:28what you get to render your scene. At a
1:30studio, you'd have to divide this time
1:32obviously with other things that want
1:34part of that time, too, because I assume
1:36your entire game isn't just about
1:38drawing trees. So, you have to share
1:39your frame budget, both CPU and GPU,
1:42with other stuff. In this little frame
1:44breakdown study on the graphics of Doom
1:46Eternal, you can see that a lot of
1:49things go into drawing a frame. You've
1:50got your GPU skinning to animate
1:52characters. There's also shadow mapping.
1:55They've got a depth prepass high zed
1:57screen space, ambient occlusion,
1:59particles, post effects, UI, whatever.
2:01There's a bunch of stuff in here. You
2:03don't have to understand it all. But
2:04during development, these would
2:06absolutely have budgets on the target
2:08hardware, how much frame time they can
2:10have in order to hit the target frame
2:12rate. Now, we need to start our
2:14investigation. Manually taking timings
2:16can get you surprisingly far. or better
2:19yet, some sort of tool to examine
2:20performance, a profiler. Depending on
2:23what you're profiling, let's
2:24oversimplify and restrict ourselves to
2:26CPU or GPU problems. Then you'll need a
2:29way of getting information about
2:30slowdowns. Some common choices would
2:33include tools by Intel, AMD, or Nvidia,
2:36stuff like Vtune, or Insight. I found a
2:39nice little list here on this Rust
2:41profiling page by Nicholas Nethercoat,
2:43who did a lot of the work with the Rust
2:45compiler performance for Firefox. We'll
2:48do some top- down analysis. If you're
2:50doing profiling, it'll generally look
2:52like this. That is starting from the
2:54top. Here's where we update all the
2:56things. That breaks up into smaller
2:58chunks where we update the foods and
3:00then we update the blah. And we'll try
3:02to see where all the time is being
3:03spent. You can see why this is called
3:05top- down analysis because you can drill
3:07down further and figure out where your
3:09overall time is being spent. bottom up
3:11by contrast. We'll look again at the
3:13frame, but we'll look at functions being
3:15used everywhere, like this matrix
3:17update. And by making those fast, you
3:19make everything fast. This is great for
3:21finding those functions where you really
3:23want to squeeze more performance out of
3:25them. In our actual case, let's figure
3:27out what's wrong. So, in the top left,
3:29I've got the frame timings running, both
3:31the CPU and GPU time, because this
3:34problem can potentially span both. We
3:36can see the CPU sucks more than the GPU.
3:38So, let's use a profiler to see why. So,
3:41we fire up a profiler of some sort. They
3:43all look different, but present similar
3:44information. Because I'm a lazy man,
3:46we're using JavaScript and Chrome's
3:48DevTool profiler, but it'll present a
3:51breakdown similar to others. So, we can
3:53see from this breakdown, we're spending
3:55well all of the time on the CPU on draw
3:58commands, which isn't exactly a
4:00surprise. We're not even drawing that
4:01many. So, we want to make this faster.
4:04How do we do that? The general steps to
4:06fix any performance problem look
4:08something like this. You establish a
4:10budget that's already done. We're
4:11dedicating our whole frame to trees. So
4:14the next steps would be well, we'll fix
4:16this problem, make it nice and small,
4:18and then if performance still sucks, we
4:20loop around and figure out what the next
4:22biggest issue is. Let's start with a
4:24really simple example. You've got this
4:26function here, really slow, and your
4:28profiler is telling you this thing is a
4:30problem. What's your first move? I mean,
4:32we could take a look at it and maybe
4:34we'll get really lucky and there's a
4:35bunch of really, really stupid things
4:37going on in there. Remove those and
4:39voila, some of our easiest performance
4:41problems are solved. If we take our tree
4:44example, well, maybe someone
4:45intentionally wrote this super poorly
4:47for YouTube video reasons. Now, we can
4:49share materials, share underlying
4:51geometries. Those begin to give you some
4:53modest gains. We've bumped up our number
4:55of trees for the same frame rate. Looks
4:57like we're hitting about 7 to 800 in
4:59this nonsense example. Even better is
5:01using GPU instancing. That's essentially
5:04telling the GPU, we're going to draw the
5:06same tree a whole bunch of times. Go do
5:08it. That gets us up to 30,000ish trees.
5:12We're also spending almost no time on
5:13the CPU now. So, we're probably
5:15bottlenecked on the GPU. Once that
5:17absurdly lowanging fruit is out of the
5:20way, now what? Well, let's back up a bit
5:22and think about things in terms of what
5:24does your system, the one you're
5:25optimizing, what does it look like at
5:28its core? Stuff goes in. These are your
5:30inputs to the system. You process them
5:33in some way and then stuff goes out. The
5:35question is then what parts of this
5:37process do you control? Because changes
5:40to this processing step here, they're at
5:42the mercy of your input. So if you
5:44control the input, it can be much more
5:46lucrative to start there instead. When
5:48you're going to change the input, often
5:50it's productive to think about, well,
5:52what does my hardware even want? What
5:55shape should the input take to maximize
5:57hardware utilization? In general,
6:00hardware likes a few things. This is a
6:02sweeping generalization, but it likes
6:04giant contiguous chunks of memory and it
6:06likes predictability. Let's say we're
6:09talking about the CPU and we're doing
6:10work with things in memory. If you
6:12process things in a contiguous chunk of
6:14memory, you process this chunk here,
6:16then afterwards you process this chunk
6:18here and then this one here, well, the
6:21CPU can boldly predict that you'll need
6:23this chunk here and just have it ready
6:25to roll when you move on to process your
6:27next chunk. Crazy, eh? But if you're
6:29randomly bouncing around memory looking
6:31for bits and pieces, the CPU has no
6:33goddamn idea where you're going to look
6:35next and so eat a big fat cache miss.
6:38Random access memory can be accessed
6:39randomly, but really shouldn't be. Cache
6:42misses are a big deal if you're not sure
6:44what they are. This little box here,
6:46let's say this is the time that it takes
6:48for the CPU to execute a simple
6:50instruction. Now, let's say that you
6:52want to do something that needs to grab
6:54something from memory. Then if it lucks
6:56out and that thing is in the L1 cache,
6:58it takes about this long in relative
7:00terms. But L1 caches are small. So if
7:03it's not there, often the CPU's next
7:06option is to check L2, which takes a bit
7:08longer, but not crazy long. Miss that
7:12though, and now you're reaching out to
7:13L3, and the cost of doing this is
7:15steadily rising. After that, if the data
7:18you're looking for isn't in any of the
7:20caches, or in other words, this is a
7:22cache miss, then you're reaching out to
7:24main memory, which takes roughly forever
7:27in computer terms. This is why cache
7:29misses hurt so much. So, if you lay out
7:31your data in a way that the CPU can
7:33predict what you're going to need next,
7:35you may eat a lot less of these. And
7:38your CPU can even try to predict which
7:40way your code is going to execute. If we
7:42compare these two loops, they're
7:44identical for the most part. So this
7:46branch here on the left predicated on
7:48data in a random order, there's no
7:50stable pattern for the CPU to pick up
7:52on. And so it's going to do so so so
7:55much worse than this one here. Even
7:57though the code itself is the same, the
7:59data is in a predictable pattern. Or
8:02since you control your inputs, you could
8:04simply separate them out into separate
8:06lists. Job done. The differences here
8:08are subtle and easy to miss. If you
8:11imagine your CPU like a little machine
8:13in a factory, you've got your programs
8:15commands coming in as a little assembly
8:17line and the results coming out as a
8:19little assembly line. Inside though, it
8:22doesn't necessarily have to be done in
8:24this order. Imagine the CPU is looking
8:26at these boxes and you can see that this
8:28box here depends on this box here and
8:31this box here, it wants something from
8:33memory. Reading from memory takes
8:35forever in CPU terms. So this whole
8:38factory operation is now blocked. The
8:41CPU wants to keep going. So maybe stuff
8:43gets shuffled a bit. We can execute this
8:45box here while we wait for the other
8:47one. Everything still comes out in the
8:49right order. Let's reset things a bit.
8:52Now imagine this box is a branch. So you
8:55got your if then else. Let's just give
8:57these different colors that you can see
8:58them clearly. Again, the CPU doesn't
9:00want to slow down, right? Figuring out
9:03this if will take a bit of time. So,
9:05while it waits to see which way it
9:07should go, it actually just tries to
9:09guess and keeps going. If they're right,
9:11awesome. If they're wrong, though, they
9:13have to backtrack and start again. It's
9:15a fantastic optimization, but
9:17mispredictions are a problem, among
9:20other things. There's so much to learn
9:22here, it's not realistic to do more than
9:24just put these on your radar. But now
9:26that you're aware, there will be links
9:28in the description, or I'll make more
9:29videos about it later. So, back to that
9:32big question from earlier. What shape
9:34should my data take? So your function
9:36here, it could be a game update loop. It
9:39could be a lot of things here. It
9:41processes a bunch of world objects which
9:43are in some random order. So we iterate,
9:46grab the entity from the world object
9:48and do our little update. From a
9:50performance standpoint, you've got
9:52pointer indirection, you've got virtual
9:54function calls, which means potential
9:55for a lot of cache misses and branch
9:57mispredictions. With some minor changes
9:59to the input, the performance of the
10:01function can substantially increase.
10:04When you're more mindful of data
10:05layouts, like in this example here, we
10:08can burn through updates way faster by
10:10packing them all into contiguous blocks
10:12of memory with no lookups. Now your CPU
10:15has one giant workload to burn through.
10:17The really cool thing here is that since
10:19your data is all tucked away nicely in a
10:21big predictable array, it gets a lot
10:23easier for either you or the compiler to
10:26vectorize this. Many CPUs have the
10:28ability to do SIMD or single instruction
10:31multiple data, meaning you can now
10:33process four or eight or even more
10:35things at once instead of one by one.
10:37The basic idea though is do your stupid
10:40things offline if possible. Once your
10:42data is in order, then you can get your
10:44code in order. Here's a few good places
10:46to start learning about this. And as a
10:48quick aside, if you've ever wondered why
10:50ECS frameworks have emerged so strongly
10:52in recent years, it should make a lot
10:54more sense now. They're in essence a
10:56data first approach. So in our tree
10:59example, we can take our tree and
11:01optimize it. Weld together redundant
11:03vertices. That'll reduce the workload.
11:05We'll reorder the mesh to maximize
11:08vertex cache usage. We'll quantize the
11:10data because you don't need these big
11:12bloated vertices. You can crunch them
11:14down quite a bit. There are a zillion
11:16tricks in this space. Ocahedron mapping
11:18for your normals. An entire tangent
11:20space can be compressed to four bytes.
11:22Positions can be done with half floats
11:24or much better 16- bit integers with
11:27scale and offset. UVs can be similarly
11:29quantized in many cases. And finally, we
11:32can compress the textures using a GPU
11:34compressed format. These might sound
11:36unimportant, but smaller vertices and
11:38textures mean better use of bandwidth
11:40and more crap fits in the cache, meaning
11:43more trees before you hit the GPU's
11:44limits. We look at our forest. It's
11:47getting bigger. You can see in this shot
11:49we've hit over 50,000 trees. So, the
11:51scene is growing quite a bit. One
11:54interesting question arises the first
11:56time you start doing these
11:57optimizations. How fast can this really
11:59go? The first couple of wins are great.
12:02The next few will probably be smaller
12:04and diminishing returns after that.
12:07Sometimes it's helpful to know what the
12:09speed of light is. Or in other words, if
12:11everything was perfect, how fast could
12:13this really go? Because you're not going
12:16to go faster than that. And as you
12:18approach that limit, the amount of time
12:20versus the amount of gain you make will
12:22really start not being worth it. So what
12:25are your other options? Do it smarter.
12:28Or said another way, don't do the things
12:30you were going to do but didn't have to.
12:33That one isn't quite as catchy. There's
12:35really nothing faster than literally
12:37just not doing it. So the first approach
12:39we talked about, we did the same work,
12:41we just made it faster to do that work.
12:44This asks the question, is there a way
12:46to literally produce the same effective
12:48output but do less overall work to get
12:50there? Can we get the same result while
12:53doing less work overall? If you're
12:55walking from A to B, well, you can get
12:57there faster by just doing an allout
12:59sprint. That's doing the same work just
13:01faster. Or you could find shortcuts.
13:04Take a little shortcut through the mall,
13:05fly over the mountain, whatever. Same
13:07destination, but you're shortening the
13:09overall route. This tends to be
13:11extremely domainspecific. So let's look
13:14at our trees. But let's look at the
13:16scene from above. If we highlight this
13:18area here, it's called the viewfrostm.
13:21It corresponds exactly to what you can
13:23see on screen. So these highlighted
13:25trees here, they're visible. While the
13:27rest of these, these are out of view. So
13:30they contribute nothing visually to the
13:31output. But in our naive setup, they
13:34still get sent to the GPU and take up
13:36processing time. So don't draw these.
13:39Simple. Call them. Bam. Instant win.
13:41You've raised the ceiling by quite a
13:42bit. How does the speed of light factor
13:44in here? Previously, it was the maximum
13:47speed that you could do certain
13:48operations like say draw trees. That
13:51hasn't really changed. If I had a
13:53theoretical max of 100 trees before,
13:55that still applies. But the difference
13:57is now they're the ones on screen and
14:00the offscreen ones don't count anymore.
14:02The ceiling still exists, but now you
14:04could have a thousand trees in the world
14:06because only 100 would be on screen at
14:08any given time. There's more avenues to
14:10explore here. In the case of our trees,
14:12we could use some sort of spatial
14:14division to rapidly discard larger
14:16portions of the world before fine grain
14:17checking each tree or take another
14:20approach. Obviously, some geological
14:22features may obstruct our view and hence
14:24anything behind that would be oluded.
14:27Watch my video on occlusion techniques
14:29to understand how much effort goes in
14:31here. But let's say that you somehow get
14:33perfect information. You're able to in
14:35the blink of an eye get a list of every
14:37single visible non-luded tree. Well,
14:40you're still bound to your speed of
14:41light restriction, unfortunately. The
14:44question then becomes, how do you break
14:45the speed of light barrier? Well,
14:47anybody who's been in a grocery store in
14:49the last couple years knows, how do you
14:51make more money without providing a
14:53better product? You don't. You cheat
14:55shamelessly and hope either nobody
14:58notices or cares enough to complain.
15:00Voila, profit. There's some really great
15:02examples over the years of games using
15:04tricks to get away with things. The fog
15:07in Silent Hill obscured your view and
15:09provided a lot of atmosphere and also
15:11ensured devs didn't have to draw too
15:13far. The infuriatingly long elevator
15:16scenes in Mass Effect gave the team time
15:18to banter, let you get to know them
15:20better, and provided cover while they
15:21loaded the scene. And Jack and Daxter
15:24streaming system would apparently have
15:26you trip and fall when you outrun the
15:28game's ability to stream in assets.
15:30That's not from any official docs, but
15:32other games have apparently done the
15:33same thing. Let's turn our attention to
15:36these trees. Here's our tree up close in
15:38all of its glory. And here it is far
15:41away, I guess, not in its glory. Well,
15:43since it's far away, why even bother
15:46drawing a good version of it? We could
15:47substitute a lesser quality version, and
15:50maybe nobody will notice or care. Maybe
15:52I already did. It's so far away. Or
15:54maybe I'm lying. Can you really tell?
15:57This is great for a variety of reasons
15:59which this brilliant YouTuber dives into
16:01in his video on why LODs or level of
16:03detail why those work so well. Beyond
16:06the savings for the smaller simplified
16:08mesh, you reduce quad overdraw. If we
16:11push it really far away, maybe we can
16:12cheat even more. Just putting a picture
16:15of a tree in. I mean, does it really
16:17matter? If this is so far away that it's
16:19being caught in the distance fog, you
16:21can't tell anyway. But the better your
16:23cheating is, the closer you can do it
16:25before anybody cares. Take this simple
16:27cube. This picture kind of sucks, but we
16:30can make it better. Maybe by separating
16:32the diffuse from the normal. Now, these
16:34can be lit nicely with some work. And we
16:36can go even further. Take a look at this
16:38technique from Ryan Brooks called
16:40octahedral impostors. The general idea
16:43is that you take that same tree and take
16:45snapshots from many different angles and
16:47use an octaedral mapping to get you a
16:49nice distribution of angles. Then with
16:51some blending, you can walk around the
16:53object and it kind of reacts like a real
16:55one. Now look at this thing. Even up
16:57close, it actually looks somewhat like
16:59the real thing, but it's all smoke and
17:02mirrors. You can scatter these out in
17:04the field now. They react to lighting.
17:06And at that point, you could even start
17:08playing with post effects like ambient
17:09occlusion or screen space shadows to
17:11make it look even better. One thing I
17:13should call out, be careful of both your
17:15average case and your worst case. If
17:18your average case is great, but every
17:20once in a while you turn the camera and
17:22hit one FPS. That's awful. As Jim kind
17:25of screamed this on Twitter, design your
17:28software to minimize fluctuations in
17:30performance. I'll give you an example.
17:33In a game, let's say you're wandering
17:34around the city. Stuff is happening blah
17:36blah, and then suddenly bombs go off,
17:38explosions, chaos. Let's focus on the
17:41particles in this case. Particles are
17:43expensive. They cause huge amounts of
17:45overdraw. So, multiple explosions go
17:47off. You flood the screen with particles
17:49and frame rate tanks. Except you can't
17:51tank the frame rate. That's not cool.
17:53You have to take this in stride. Well,
17:55you know roughly how fast you can draw
17:57particles, right? You've done your
17:58tests. Speed of light, yada yada. On
18:01prototype back in 2005, 2006, I did a
18:04dumb but surprisingly effective Hail
18:06Mary change. Moments before a big
18:09publisher demo, I put a hacky throttle
18:11on the effects. When all those
18:13explosions went off, we didn't just
18:15spawn everything. We pulled from a
18:17shared budget and damped how many
18:19particles each effect could actually
18:20use. Scaling the density of effects.
18:23This evolved into a game of how do we
18:25distribute the particle budget? Closer,
18:28more important stuff, more budget.
18:30Further away things, less one explosion
18:33or 100, things should stay even. Newer
18:35techniques are much smarter than that,
18:37doing things like dynamically adjusting
18:39the resolution. But the idea is just
18:41don't let the worst case be much
18:43different than your average case. So
18:45overall, that's kind of how optimization
18:48works. We started with our simple scene
18:50unoptimized, and we worked through the
18:52various stages. We got the really
18:54lowhanging fruit out of the way,
18:56improved our input, and made sure to use
18:58the hardware's preferred way of drawing
19:00lots of stuff and saw the scene's
19:02performance grow. After that, we got a
19:04bit smarter about distributing the work,
19:06calling out what was invisible, dividing
19:08the world spatially into regions to
19:10rapidly test. And again, we saw big
19:12improvements. Eventually, you hit
19:14diminishing returns because your
19:16hardware has real verifiable limits. And
19:18that's when you start trying to get
19:20creative. What will people notice? What
19:22can I get away with doing? Do the same
19:24work but better. Do less work. Be
19:26smarter. Cheat and lie. In reality,
19:29optimization work is a mixed bag of
19:30these approaches. But let's not nitpick.
19:33It's generally always that same
19:35approach. You figure out how much you
19:37want to spend on this. You profile and
19:39figure out where your problems are. Find
19:40a problematic bit of code. Make changes
19:43there and profile again. The details
19:45will change as you get more experience.
19:47But let's not over complicate this. If
19:49you came into this video thinking I'd
19:50make you some sort of optimization
19:52monster in 20 minutes or however long
19:54this took, sorry. I've seen even
19:57experienced game devs, guys with a
19:59decade or more experience, blindly
20:02optimizing with no real plan. So, just
20:04getting a good framework as a start. I
20:06don't pretend to be the best or to know
20:08everything. For any devs who have some
20:10interesting war stories of major
20:12performance improvements, comment
20:13sections below. Make sure to share.
20:16Cheers.