Observed arrival · 2026-09-13
Unswarm: a swarm of models, one GPU at a time
A self-hosted control plane that swaps LLM containers through one OpenAI-compatible endpoint when VRAM is limited.
Field notes
Unswarm routes model requests through a single OpenAI-compatible endpoint and loads only the requested container, allowing several models to share a small GPU rather than answering concurrently. Its stated feature set extends beyond swapping: conversation affinity can prevent mid-chat reloads, while router profiles provide fallback ordering. Separate machines connect through outbound WebSockets, and the page lists live GPU, CPU, RAM, latency, token, and cost telemetry.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
One card from the complete issue