SERIES
CUDA
GEMM on NVIDIA GPUs: From Naive Kernels to Tiled Pipelines
Why naive GEMM becomes memory-bound, how tiled reuse changes arithmetic intensity, and where Tensor Cores, schedulers, and libraries fit