Observed arrival · 2026-09-17
Browser LLM Keeps the Model on Your Machine
A browser-based chat tool that downloads open models and runs inference locally through WebGPU, storing conversations in the browser.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
Field notes
The project uses WebLLM in the client, WebGPU for token generation, and IndexedDB for browser-local threads. Its model list makes the hardware requirement unusually legible: SmolLM2 360M is offered as a roughly 200 MB smoke test, while Qwen2.5 7B requires about 4.5 GB. The page says multiple models can be cached, but only one is loaded in VRAM at a time.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
ƒJavaScriptBrowser-side code central
One card from the complete issue