Observed arrival · 2026-09-14
InferenceGuru: Inside the Inference Machine
A hands-on course explaining how models run, from laptop-scale predictions to GPU fleets serving language, vision, speech, and recommendation systems.
Field notes
The course begins with a model file and a prediction on a local laptop before moving through GPU memory, kernels, runtimes, token-by-token language generation, and fleet-scale serving. Its LLM-engine section proposes starting vLLM, SGLang, and TensorRT-LLM with the same script for comparison. The homepage says eight lessons are ready, while many later topics remain marked “coming,” making the current boundary of the material explicit.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue