Observed arrival · 2026-09-23
EcoCompute measures the energy cost of LLM quantization
An open benchmark protocol for measuring how LLM quantization affects GPU energy use, with data, runnable tools, and a public coverage matrix.
- For
- Researchers and engineers comparing LLM inference energy
- Worth noticing
- Both the container and free Colab notebook emit the same energy.json format; contributors preview results in-browser before opening a GitHub issue.
Field notes
The protocol fixes decoding at batch size one and 256 tokens, samples NVML power at 10 Hz, and limits its figures to GPU-package draw rather than whole-system energy. A free Colab notebook and an MLCube-compatible container produce the same schema-validated file, which contributors can inspect against the published curve before preparing a GitHub issue. The page also distinguishes its main dataset's reported n=2 from a deeper RTX 4090 set with n=1 per configuration.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue