Observed arrival · 2026-08-21
Throttle measures the price of idle GPUs
Throttle is a drop-in proxy for self-hosted LLM inference that promises lower costs through dynamic batching, instance right-sizing, spot-aware capacity, and semantic caching.
Why it surfaced
This pre-launch project makes a concrete case instead of waving at AI efficiency: its homepage reports 991 tok/s versus 312 tok/s with batching enabled, a 3.18× throughput increase across 1,206 requests. It also links to source code and invites teams to submit their engine, model, and traffic shape for hands-on onboarding.
Throttle is a drop-in proxy for self-hosted LLM inference that promises lower costs through dynamic batching, instance right-sizing, spot-aware capacity, and semantic caching.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue