I’ve been working on a renderer for my custom engine using the GPU API, but I’ve noticed that it performs considerably worse than the fallback renderer I made using OpenGL.
In trying to diagnose the issue I created a new heavily simplified example program from scratch, both with OpenGL and SDL GPU.
- In the first iteration I set up an index/vertex buffers and uploaded data to them once, then rendered the triangles each frame. OpenGL and SDL GPU performance is the same.
- Next I changed things so that I’m uploading data to the buffers each frame, before rendering it (I’m replacing the whole buffers each time). OpenGL performance is indistinguishable from the previous case, while SDL GPU now is 50% slower (in terms of framerate in immediate mode). My implementation uses cycling when uploading data to the buffers, but toggling cycling on/off does not seem to impact performance.
I must be doing something wrong here but I can’t figure out what. Does anyone with more experience with the API have any idea how to do this properly?
The code that I’m using in the draw call each frame is attached below. If it helps, I’m testing this on Linux with the Vulkan backend, on an AMD Radeon RX 6650 XT with the open source Vulkan driver.
void SDLTest::draw() {
cmdBuffer = SDL_AcquireGPUCommandBuffer(device);
SDL_WaitAndAcquireGPUSwapchainTexture(cmdBuffer, window, &swapchainTexture, nullptr, nullptr);
if (!swapchainTexture) {
SDL_SubmitGPUCommandBuffer(cmdBuffer);
return;
}
// copy data to vertex and index buffers
auto copyPass = SDL_BeginGPUCopyPass(cmdBuffer);
{
size_t size = sizeof(Vertex) * vertices.size();
auto data = SDL_MapGPUTransferBuffer(device, transferBuffers.vertex, true);
memcpy(data, vertices.data(), size);
SDL_UnmapGPUTransferBuffer(device, transferBuffers.vertex);
SDL_GPUTransferBufferLocation location{};
location.transfer_buffer = transferBuffers.vertex;
location.offset = 0;
SDL_GPUBufferRegion region{};
region.buffer = buffers.vertex;
region.size = size;
region.offset = 0;
SDL_UploadToGPUBuffer(copyPass, &location, ®ion, true);
}
{
size_t size = sizeof(uint16_t) * indices.size();
auto data = SDL_MapGPUTransferBuffer(device, transferBuffers.index, true);
memcpy(data, indices.data(), size);
SDL_UnmapGPUTransferBuffer(device, transferBuffers.index);
SDL_GPUTransferBufferLocation location{};
location.transfer_buffer = transferBuffers.index;
location.offset = 0;
SDL_GPUBufferRegion region{};
region.buffer = buffers.index;
region.size = size;
region.offset = 0;
SDL_UploadToGPUBuffer(copyPass, &location, ®ion, true);
}
SDL_EndGPUCopyPass(copyPass);
SDL_GPURenderPass* renderPass;
{
SDL_GPUColorTargetInfo colorTargetInfo{};
colorTargetInfo.load_op = SDL_GPU_LOADOP_CLEAR;
colorTargetInfo.clear_color = {.5, .5, .5, 1.};
colorTargetInfo.store_op = SDL_GPU_STOREOP_STORE;
colorTargetInfo.texture = swapchainTexture;
renderPass = SDL_BeginGPURenderPass(cmdBuffer, &colorTargetInfo, 1, nullptr);
}
// bind the vertex buffer
{
SDL_GPUBufferBinding bufferBindings {};
bufferBindings.buffer = buffers.vertex;
bufferBindings.offset = 0;
SDL_BindGPUVertexBuffers(renderPass, 0, &bufferBindings, 1);
}
// bind the index buffer
{
SDL_GPUBufferBinding bufferBindings {};
bufferBindings.buffer = buffers.index;
bufferBindings.offset = 0;
SDL_BindGPUIndexBuffer(renderPass, &bufferBindings, SDL_GPU_INDEXELEMENTSIZE_16BIT);
}
SDL_BindGPUGraphicsPipeline(renderPass, pipeline);
SDL_DrawGPUIndexedPrimitives(renderPass, indices.size(), 1, 0, 0, 0);
SDL_EndGPURenderPass(renderPass);
SDL_SubmitGPUCommandBuffer(cmdBuffer);
}