Skip to the card

Card 912 of 9842026-09-13 issue

Observed arrival · 2026-09-13

Unswarm: a swarm of models, one GPU at a time

unswarm.dev Observed source
Editorial interest 86/100 Selection signal · not a rating of the site

A self-hosted control plane that swaps LLM containers through one OpenAI-compatible endpoint when VRAM is limited.

Landing page captured for the 2026-09-13 issue.

Field notes

Unswarm routes model requests through a single OpenAI-compatible endpoint and loads only the requested container, allowing several models to share a small GPU rather than answering concurrently. Its stated feature set extends beyond swapping: conversation affinity can prevent mid-chat reloads, while router profiles provide fallback ordering. Separate machines connect through outbound WebSockets, and the page lists live GPU, CPU, RAM, latency, token, and cost telemetry.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

OpenPublic substance visible
LoginAccess appeared gated
PrettyNotable craft visible
ProPolished or operationally mature
NicheUnusually specific use

One card from the complete issue

W Has No Surviving Value

304,798 arrived 1,000 judged 984 catalogued Enter the complete issue
unswarm.dev

Landing page observed 2026-09-13. The live site may have changed.