Observed arrival · 2026-09-20
Roofline Labs Measures the Cost of Making LLMs Move
Roofline Labs benchmarks self-hosted LLM inference against hardware memory-bandwidth limits, tracking utilization, tokens per second, and cost per token.
- For
- Self-hosted LLM infrastructure teams
- Worth noticing
- The public table records 45 decode results on an NVIDIA DGX Spark, scored against the hardware roofline.
Field notes
The page grounds its central idea in a memory-bandwidth ceiling: a decode-bound model cannot exceed the available bandwidth divided by bytes read per token. Its displayed GB10 comparisons include Qwen3.6-35B-A3B under multiple engines and speculation settings, plus a substantially lower result for Qwen3-Coder-Next 80B-A3B. The broader fleet-monitoring and recommendation workflow is presented as the company’s stated product direction.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue