Observed arrival · 2026-08-25
Learn Kernels, a field guide to GPU performance
An interactive book explains GPU kernels, inference performance, optimization, profiling, and distributed serving.
Field notes
The book arranges GPU performance as a hierarchy, beginning with threads, blocks, and warps before moving through kernel optimization, inference engines, and distributed serving. Its scope includes matrix multiplication, tensor cores, attention, quantization, continuous batching, mixture-of-experts systems, and current hardware such as NVIDIA Blackwell, AMD, TPU, and Trainium. The homepage also exposes a glossary, reading list, developer area, API, MCP, and llms.txt resources.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue