Observed arrival · 2026-09-22
Fusion Runtime keeps the whole voice conversation on one GPU
A self-hosted voice-agent runtime that runs speech recognition, an LLM, and text-to-speech in one process on hardware you control.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Developers building private, low-latency voice interfaces
- Worth noticing
- The published benchmark separates roughly 490 ms of processing from a roughly 500 ms configurable silence wait.
Field notes
The documented pipeline combines Silero VAD, faster-whisper, llama.cpp, and Kokoro without a network hop between stages. The browser client handles echo cancellation and microphone resampling, while the backend mints a one-use token rather than exposing an API key in the page. Reported measurements use Qwen 7B, Whisper small, and Kokoro on an RTX 3090; long-call behavior remains unmeasured.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue