Skip to the card

Card 541 of 9752026-09-24 issue

Observed arrival · 2026-09-24

MLX Peer: splitting local inference across a Mac and iPhone

mlx-peer.com Visit website
Editorial interest 83/100 Selection signal · not a rating of the site

An experimental project explores dividing language-model inference between a Mac and iPhone connected over USB.

Landing page captured for the 2026-09-24 issue.
For
Developers experimenting with Apple Silicon local inference
Worth noticing
The 27B run used 12 of 64 layers on iPhone; its prefill comparison missed the stated numerical tolerance.

Field notes

The page describes a layer-partitioned inference pass: the Mac handles embeddings and output, while the iPhone runs layers 0–11 and returns activations over USB. Its exporter writes separate weight files without first allocating the full model in memory. In the cited 27B run, the phone held 2.57 GB of weights and decode reached 3.93 tokens per second after the first token; the page flags substantial Mac swap and a failed prefill tolerance check.

Observed signals

Read the marks

Editorial observations of this landing page, not a rating.

○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use

One card from the complete issue

The Busker’s Compute Bill

314,149 arrived 1,000 judged 975 catalogued Enter the complete issue
mlx-peer.com

Landing page observed 2026-09-24. The live site may have changed.