A real-time 3D Gaussian Splatting rasterizer written from scratch in CUDA, with an OpenGL window for presentation. It loads a trained .ply capture (the format the INRIA training code exports) and flies a camera through it. No splatting library is used. The projection, tiling, sorting and compositing all live in cuda_render.cu; OpenGL only owns the window and blits the CUDA-written texture to the screen.

The renderer running on the built-in demo scene

How it renders a frame

Splatting is order-dependent alpha compositing over millions of primitives, so the whole design is about getting them into per-pixel depth order cheaply. The frame runs as four GPU passes.

1. Preprocess

preprocess_splats projects every splat once. The centre goes through a pinhole projection; the 3D covariance R S Sᵀ Rᵀ is pushed through a local affine approximation of that projection (the EWA method) to get a 2D covariance, which is inverted into a conic so the render pass can evaluate a Gaussian with three multiplies. The larger eigenvalue of the 2D covariance is the variance along the splat's long axis, so 3σ of it bounds the on-screen footprint in every direction. That bound is clipped to the 16×16 tile grid, and the pass records how many tiles the splat covers.

2. Duplicate

duplicate_splats expands each splat into one entry per tile it touches. A prefix sum over the per-splat tile counts gives each splat a private slice of the output array, so the expansion needs no atomics. Each entry gets a 64-bit key: the tile id in the high 32 bits, the float depth's raw bits in the low 32. Positive floats compare correctly as integers, so one radix sort over that key groups entries by tile and orders them by depth within each tile at the same time. Only the bits the tile id actually occupies are sorted above the depth bits.

3. Tile ranges

identify_tile_ranges walks the sorted keys and records where each tile's run of entries starts and ends.

4. Composite

render_tiles runs one block per tile and one thread per pixel. The block cooperatively loads its tile's splats into shared memory in 256-entry batches and each thread walks them front to back, accumulating colour and attenuating transmittance. A thread that saturates stops accumulating but keeps reaching the barriers, so the early exit is a block-wide __syncthreads_count rather than a per-thread break. The result is written straight into an OpenGL texture through CUDA/GL interop, so the pixels never travel back through host memory.