GPU API: performance issues when uploading vertex/index buffer each frame

I’ve been working on a renderer for my custom engine using the GPU API, but I’ve noticed that it performs considerably worse than the fallback renderer I made using OpenGL.

In trying to diagnose the issue I created a new heavily simplified example program from scratch, both with OpenGL and SDL GPU.

  • In the first iteration I set up an index/vertex buffers and uploaded data to them once, then rendered the triangles each frame. OpenGL and SDL GPU performance is the same.
  • Next I changed things so that I’m uploading data to the buffers each frame, before rendering it (I’m replacing the whole buffers each time). OpenGL performance is indistinguishable from the previous case, while SDL GPU now is 50% slower (in terms of framerate in immediate mode). My implementation uses cycling when uploading data to the buffers, but toggling cycling on/off does not seem to impact performance.

I must be doing something wrong here but I can’t figure out what. Does anyone with more experience with the API have any idea how to do this properly?

The code that I’m using in the draw call each frame is attached below. If it helps, I’m testing this on Linux with the Vulkan backend, on an AMD Radeon RX 6650 XT with the open source Vulkan driver.

void SDLTest::draw() {
    cmdBuffer = SDL_AcquireGPUCommandBuffer(device);

    SDL_WaitAndAcquireGPUSwapchainTexture(cmdBuffer, window, &swapchainTexture, nullptr, nullptr);
    if (!swapchainTexture) {
        SDL_SubmitGPUCommandBuffer(cmdBuffer);
        return;
    }

    // copy data to vertex and index buffers
    auto copyPass = SDL_BeginGPUCopyPass(cmdBuffer);

    {
        size_t size = sizeof(Vertex) * vertices.size();
        auto data = SDL_MapGPUTransferBuffer(device, transferBuffers.vertex, true);
        memcpy(data, vertices.data(), size);
        SDL_UnmapGPUTransferBuffer(device, transferBuffers.vertex);

        SDL_GPUTransferBufferLocation location{};
        location.transfer_buffer = transferBuffers.vertex;
        location.offset = 0;

        SDL_GPUBufferRegion region{};
        region.buffer = buffers.vertex;
        region.size = size;
        region.offset = 0;

        SDL_UploadToGPUBuffer(copyPass, &location, &region, true);
    }

    {
        size_t size = sizeof(uint16_t) * indices.size();
        auto data = SDL_MapGPUTransferBuffer(device, transferBuffers.index, true);
        memcpy(data, indices.data(), size);
        SDL_UnmapGPUTransferBuffer(device, transferBuffers.index);

        SDL_GPUTransferBufferLocation location{};
        location.transfer_buffer = transferBuffers.index;
        location.offset = 0;

        SDL_GPUBufferRegion region{};
        region.buffer = buffers.index;
        region.size = size;
        region.offset = 0;

        SDL_UploadToGPUBuffer(copyPass, &location, &region, true);
    }

    SDL_EndGPUCopyPass(copyPass);

    SDL_GPURenderPass* renderPass;

    {
        SDL_GPUColorTargetInfo colorTargetInfo{};
        colorTargetInfo.load_op = SDL_GPU_LOADOP_CLEAR;
        colorTargetInfo.clear_color = {.5, .5, .5, 1.};
        colorTargetInfo.store_op = SDL_GPU_STOREOP_STORE;
        colorTargetInfo.texture = swapchainTexture;

        renderPass = SDL_BeginGPURenderPass(cmdBuffer, &colorTargetInfo, 1, nullptr);
    }

    // bind the vertex buffer
    {
        SDL_GPUBufferBinding bufferBindings {};
        bufferBindings.buffer = buffers.vertex;
        bufferBindings.offset = 0;
        SDL_BindGPUVertexBuffers(renderPass, 0, &bufferBindings, 1);
    }

    // bind the index buffer
    {
        SDL_GPUBufferBinding bufferBindings {};
        bufferBindings.buffer = buffers.index;
        bufferBindings.offset = 0;
        SDL_BindGPUIndexBuffer(renderPass, &bufferBindings, SDL_GPU_INDEXELEMENTSIZE_16BIT);
    }

    SDL_BindGPUGraphicsPipeline(renderPass, pipeline);

    SDL_DrawGPUIndexedPrimitives(renderPass, indices.size(), 1, 0, 0, 0);

    SDL_EndGPURenderPass(renderPass);

    SDL_SubmitGPUCommandBuffer(cmdBuffer);
}

I don’t have experience in this part of SDL, but can you check off each member in this list from SDL’s 3D API doc

Performance considerations

Here are some basic tips for maximizing your rendering performance.

  • Beginning a new render pass is relatively expensive. Use as few render passes as you can.
  • Minimize the amount of state changes. For example, binding a pipeline is relatively cheap, but doing it hundreds of times when you don’t need to will slow the performance significantly.
  • Perform your data uploads as early as possible in the frame.
  • Don’t churn resources. Creating and releasing resources is expensive. It’s better to create what you need up front and cache it.
  • Don’t use uniform buffers for large amounts of data (more than a matrix or so). Use a storage buffer instead.
  • Use cycling correctly. There is a detailed explanation of cycling further below.
  • Use culling techniques to minimize pixel writes. The less writing the GPU has to do the better. Culling can be a very advanced topic but even simple culling techniques can boost performance significantly.

In general try to remember the golden rule of performance: doing things is more expensive than not doing things. Don’t Touch The Driver!

Unfortunately nothing in the list stands out. I made this example specifically to have the smallest number of moving parts: there’s only one render pass, no new resource are created, no state changes within each frame.

The only difference with all the standard examples of how to use the GPU API is that I’m reuploading the data to the buffers each frame before drawing to the screen.

Some more information following a few more tests: I tried removing the render pass completely, doing only the copy pass each frame, and the performance stays effectively the same (slightly faster because no drawing happens, but still roughly half of the performance when the buffers are uploaded once).

I also tried changing the number of frames in flight: 3 does not improve performance, 1 makes it only slightly slower.

Here are a couple of tutorials, but on further reading below, I don’t think there is actually a problem to be fixed.

I’m pretty sure the underlying gpu driver on most systems will be Vulkan. One of the goals of SDL3’s API is that you can swap it out with D3D, or Metal, or other available drivers with minimal code change on your part.

Unfortunately, I see here that the default underlying driver, Vulkan, is not built to be faster than OpenGL at simple operations. That was never the goal. It instead utilizes less CPU, manages multithreading better, and gives the developer more options.

So the SDL3 GPU API likely would be an improvement over OpenGL when your game starts to get CPU bound, when the scenes hit increased complexity, and when you feel like the capabilities of OpenGL API have become the limitation (though as a wrapper, SDL3 may introduce their own limitations in this area).

If you are more comfortable using OpenGL and don’t foresee any of the above situations for the project, then it is perfectly acceptable to skip the SDL3 GPU API and use OpenGL for your game.

I am aware that Vulkan only provides gains in specific situations. I’m not expecting SDL GPU to perform better than OpenGL in this simple example, the performance issue I’m noticing is that uploading the buffer data each frame makes the frame rate drop in half when using the GPU API and it leaves it the same with OpenGL. That’s what’s leading me to think I’m doing something wrong.

I took a look at the tutorials but unfortunately, like all of the examples from the github repo, the buffers are only uploaded once at the beginning.

I had to ask AI since I lack personal experience/research.
Here is what google says:

Mapping Transfer Buffers Inside a Copy Pass

You are calling SDL_MapGPUTransferBuffer and SDL_UnmapGPUTransferBuffer inside an active copy pass (between SDL_BeginGPUCopyPass and SDL_EndGPUCopyPass). [1, 2]

  • The Rule: Mapping and unmapping occur on the CPU timeline immediately, whereas a copy pass records commands for the GPU timeline. All CPU-side modifications to transfer buffers should be completed and unmapped before you open a pass to encode commands targeting those buffers. [1, 2, 3]

Sorry, but I hope that helps.

That’s not it. In fact, removing the SDL_MapGPUTransferBuffer/memcpy/SDL_UnmapGPUTransferBuffer step completely (doing it only once to populate the data) and leaving only the SDL_UploadToGPUBuffer to be done each frame does not improve performance at all.

I should check, you are saying that you are getting half the FPS with SDL_GPU than what you get with OpenGL, but what are the actual FPS you are seeing? 10, 30, 50, 120, 3000?

If we are talking SDL getting 4000 FPS while OpenGL is at 8000, that actually does not mean very much, that is easily just the result of different overheads. The “slower one” may actually be setting up for heavy work while the other skips the homework. However, if we were talking even 55 FPS on SDL vs 60 with OpenGL, I would consider that to be much more critical.

Further checking with AI:

The behavior the OP is describing—a massive performance drop when updating buffers every frame in a modern API (SDL3 GPU) that doesn’t happen in an older API (OpenGL)—is a textbook example of a GPU pipeline stall.

In modern, low-level graphics APIs, the driver no longer hides synchronization for you. If you overwrite a buffer that the GPU is still reading from to render the previous frame, the hardware must stop everything until that read is finished. The fact that the OP tried changing “frames in flight” without success strongly suggests they only increased the number of transfer buffers, while still targeting a single destination buffer. This confirms the bottleneck is the destination synchronization. Implementing a ring buffer for the destination buffers is the industry-standard solution for this specific problem.

By destination buffers I’m pretty sure it means both the vertex and index buffers.

An SDL_GPUBuffer object is already a ring buffer behind the scenes (and I’m using cycling when uploading data, so I’m swapping buffers already) so that’s also incorrect.

I appreciate that you’re trying to help, I really do, but I’d rather not be subjected to second-hand AI answers (If I wanted those—and I most definitely don’t—I could have gotten them myself). It’s ok not to know the answer to something. I’m not in a rush, I’m happy to wait for someone familiar with the API to provide some insight.

3 Likes

Fair enough, I wish I could have done more. Let me try to catch some attention.
Activating Bat Signal: (I hope this isn’t rude)
@cosmonaut , @icculus, @TheSpydog; Hi, I hope you are doing well.
The 3D API is way beyond me, I’m totally lost in this topic. This thread has 158 views so far and no other help has come through to save the day. I’m hoping you might have a minute to lift the fog.

1 Like

Thanks, I really appreciate that. I hope I didn’t come off as rude earlier.

I should check, you are saying that you are getting half the FPS with SDL_GPU than what you get with OpenGL, but what are the actual FPS you are seeing? 10, 30, 50, 120, 3000?

This is a critical observation. Benchmarking trivial workflows is a noob trap. If you are seeing significant performance loss with an actual workflow that would be of concern, but benchmarking a simple workflow doesn’t tell us anything - Vulkan has overhead that pays off only when there is more work to do.

EDIT: Also worth noting that at large enough scales, an FPS drop can be trivial in terms of time taken (for example, a drop from 8000FPS to 4000FPS is a difference of 0.1ms). Frame time is generally a better assessment of how long things are taking.

2 Likes

I should have made this clearer in the original post: I did see performance loss in an actual workflow (the renderer I wrote for a custom game engine) and that is what prompted this question; on the low-end machine I tested it on I could get to 100 FPS for OpenGL and 50 FPS for SDL_GPU given enough triangles to draw. To investigate the issue I made this trivial example that only keeps the main feature of my renderer, i.e. the vertex/index buffers updated each frame and used in a single draw call, that I use for automatic sprite batching.

I am not actually trying to benchmark the trivial example. The issue (which I believe is the cause of the performance loss in my renderer) is that the SDL_GPU implementation stalls between draw calls by waiting for the previous draw operation to finish before copying the new data to the buffers, despite the fact that I’m using cycling both for mapping the transfer buffers and copying to the actual buffers. The OpenGL implementation doesn’t.

Indeed, I just confirmed that I can make OpenGL have the same performance loss by not using GL_MAP_INVALIDATE_BUFFER_BIT when writing the data to the buffers, which forces it to wait for the previous draw call to finish.

Update on the issue: I rebuilt the example directly in Vulkan to figure out what’s going wrong.

Initially I replicated the same setup (transfer buffers on coherent memory and the vertex/index buffers on device local memory, all on ring buffers to avoid stalling) and observed the exact same performance that SDL_GPU has.

However, I was eventually able to get the same performance that my OpenGL implementation has, by skipping the transfer buffers completely and allocating the vertex/index buffers on coherent memory and writing the data directly to them with memcpy each frame. I guess that when uploading data every frame the benefits of using device local memory are negligible compared to the cost of copying from the transfer buffers.

Now, getting back to SDL_GPU: as far as I understand there is no equivalent of using coherent, host visible memory for the vertex/index buffers—they’re always allocated as device local and require a transfer buffer to send data to them. Is this correct? In this case it looks like there’s no way to avoid this performance loss.

If you don’t get a swapchain texture from SDL_WaitAndAcquireGPUSwapchainTexture() then you should call SDL_CancelGPUCommandBuffer() instead of submitting an empty command buffer.

I made an app with SDL_GPU a while back, and while it didn’t upload new vertex data every frame, it used Dear ImGUI (which does upload new vertex data every frame, for thousands of vertexes) and it didn’t have any performance issues. I’m noticing that their SDL_GPU backend handles the transfer buffers outside of the copy pass.

Perhaps take a peek at Dear ImGUI’s SDL_GPU backend, or SDL_Renderer’s own SDL_GPU backend.

I did take a look at the Dear ImGUI SDL_GPU backend when I was trying to figure things out, and it uses pretty much the same approach I’m using. Writing the data to the transfer buffer inside or outside the copy pass does not make a difference (which makes sense, the copy pass is just being recorded in the command buffer, not being exectuted right at that moment). Now that I think of it, when doing the testing of the renderer in my full engine (which also uses Dear ImGUI) I noticed the performance difference OpenGL vs SDL_GPU even when only drawing ImGUI stuff, so not touching my buffers at all.

I hadn’t thought of checking the SDL_Renderer source code, just took a look at it now. It also does the same thing I’m doing.