A real-time 3D Gaussian Splatting rasterizer written from scratch in CUDA, with an OpenGL
window for presentation. It loads a trained .ply capture (the format the INRIA
training code exports) and flies a camera through it. No splatting library is used. The
projection, tiling, sorting and compositing all live in cuda_render.cu; OpenGL
only owns the window and blits the CUDA-written texture to the screen.
How it renders a frame
Splatting is order-dependent alpha compositing over millions of primitives, so the whole design is about getting them into per-pixel depth order cheaply. The frame runs as four GPU passes.
1. Preprocess
preprocess_splats projects every splat once. The centre goes through a pinhole
projection; the 3D covariance R S Sᵀ Rᵀ is pushed through a local
affine approximation of that projection (the EWA method) to get a 2D covariance, which is
inverted into a conic so the render pass can evaluate a Gaussian with three multiplies. The
larger eigenvalue of the 2D covariance is the variance along the splat's long axis, so
3σ of it bounds the on-screen footprint in every direction. That bound is clipped to the
16×16 tile grid, and the pass records how many tiles the splat covers.
2. Duplicate
duplicate_splats expands each splat into one entry per tile it touches. A prefix
sum over the per-splat tile counts gives each splat a private slice of the output array, so
the expansion needs no atomics. Each entry gets a 64-bit key: the tile id in the high 32 bits,
the float depth's raw bits in the low 32. Positive floats compare correctly as integers, so
one radix sort over that key groups entries by tile and orders them by depth within
each tile at the same time. Only the bits the tile id actually occupies are sorted above the
depth bits.
3. Tile ranges
identify_tile_ranges walks the sorted keys and records where each tile's run of
entries starts and ends.
4. Composite
render_tiles runs one block per tile and one thread per pixel. The block
cooperatively loads its tile's splats into shared memory in 256-entry batches and each thread
walks them front to back, accumulating colour and attenuating transmittance. A thread that
saturates stops accumulating but keeps reaching the barriers, so the early exit is a
block-wide __syncthreads_count rather than a per-thread break. The
result is written straight into an OpenGL texture through CUDA/GL interop, so the pixels never
travel back through host memory.