HomeAll ToolsCategories
Browse Tools

backburner

Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

free(Free tier)848 github stars

About backburner

Plug your iPhone into your MacBook with a 10 Gb/s USB-C cable and it helps run Qwen3.8-27B locally:

Faster prefill (up to 64k context). For every batch of prompt tokens the Mac runs layers 1-40 and the iPhone runs 41-64 on its GPU, pipelined. Your agent waits less every time it reads a file or a tool result of more than ~512 tokens: 29-44% faster prefill at 16k-48k. More context. A 24 GB Mac fits 64k tokens of 8-bit context next to the model. The iPhone holds the oldest part past that and computes attention over it: its GPU during prefill, its GPU and Neural Engine while writing.

The engine is a llama.cpp fork (llama.cpp/, StayLameBro/backburner-llama.cpp) with its own Mac kernels (SME2, Metal fusions, DFlash2 speculative decoding). Those speed things up on the Mac alone too; the numbers below keep the two apart.