Observed arrival · 2026-09-03
Uno GPU wants private models to sleep between thoughts
Uno GPU presents dedicated GPU slices that wake in 0.6 seconds, run a user’s model in a Linux VM, and bill only while the machine is thinking.
Field notes
Uno GPU describes a snapshot-based inference workflow: a model state sits beside the GPU, streams into VRAM on request, serves through an OpenAI-compatible endpoint, and returns to sleep when idle. Its worked example uses Qwen3 27B, a 27 GB state, and an H200 slice with 35 GB of VRAM; the page states a 0.58-second resume and 0.60-second first-token path. The GPU offering is presented as early access, while the underlying snapshot mechanism is attributed to Uno Cloud production use.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue