Observed arrival · 2026-09-14
Horizon RunMap — MoE architecture by Lucas Ribeiro
An independent technical preview explaining how mixture-of-experts models can keep compact expert weights in RAM while moving selected computations to the GPU.
Field notes
The architecture keeps the complete immutable expert-weight source in RAM and uses GPU memory as a working area for the experts selected during inference. Its GLM-5.3 example contrasts roughly 1.49 TB of BF16 weight memory with 744 GB in FP8, then relates those figures to B200 capacity, server headroom, context length, and concurrency. The page also exposes audit material, JSON data, and links to model specifications and deployment references.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue