Observed arrival · 2026-09-24
MLX Peer: splitting local inference across a Mac and iPhone
An experimental project explores dividing language-model inference between a Mac and iPhone connected over USB.
○Open
⊠Login
$Paid
†Ads
✦Pretty
●Pro
◎Niche
◉Human
⚑Risk
ƒJS
- For
- Developers experimenting with Apple Silicon local inference
- Worth noticing
- The 27B run used 12 of 64 layers on iPhone; its prefill comparison missed the stated numerical tolerance.
Field notes
The page describes a layer-partitioned inference pass: the Mac handles embeddings and output, while the iPhone runs layers 0–11 and returns activations over USB. Its exporter writes separate weight files without first allocating the full model in memory. In the cited 27B run, the phone held 2.57 GB of weights and decode reached 3.93 tokens per second after the first token; the page flags substantial Mac swap and a failed prefill tolerance check.
Observed signals
Read the marks
Editorial observations of this landing page, not a rating.
○OpenPublic substance visible
✦PrettyNotable craft visible
●ProPolished or operationally mature
◎NicheUnusually specific use
One card from the complete issue