Skip to the card

Card 715 of 9932026-09-20 issue

Observed arrival · 2026-09-20

Roofline Labs Measures the Cost of Making LLMs Move

rooflinelabs.com Visit website
Editorial interest 84/100 Selection signal · not a rating of the site

Roofline Labs benchmarks self-hosted LLM inference against hardware memory-bandwidth limits, tracking utilization, tokens per second, and cost per token.

Landing page captured for the 2026-09-20 issue.
For
Self-hosted LLM infrastructure teams
Worth noticing
The public table records 45 decode results on an NVIDIA DGX Spark, scored against the hardware roofline.

Field notes

The page grounds its central idea in a memory-bandwidth ceiling: a decode-bound model cannot exceed the available bandwidth divided by bytes read per token. Its displayed GB10 comparisons include Qwen3.6-35B-A3B under multiple engines and speculation settings, plus a substantially lower result for Qwen3-Coder-Next 80B-A3B. The broader fleet-monitoring and recommendation workflow is presented as the company’s stated product direction.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

Mahjongg With Footnotes

304,744 arrived 1,000 judged 993 catalogued Enter the complete issue
rooflinelabs.com

Landing page observed 2026-09-20. The live site may have changed.