StayLameBro/backburner

Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

Open original source ↗

From the publisher

README & documentation

Source preview

Read the project’s overview, installation instructions and usage examples. The original README is the source of truth.

Plug your iPhone into your MacBook with a 10 Gb/s USB-C cable and it helps run Qwen3.8-27B locally: on its GPU, pipelined. Your agent waits less every time it reads a file or a tool result: 29-44% faster prefill at 16k-48k. and computes attention over it: its GPU during prefill, its GPU and Neural Engine while writing. The server sizes the total from the phone's free memory at startup (196k-229k tokens at 8-bit on an iPhone 17 Pro Max). Tested end to end to 128k at 8-bit and 140k at 4-bit. The engine is a llama.cpp fork (llama.cpp/, StayLameBro/backburner-llama.cpp) with its own Mac kernels (SME2, Metal fusions, DFlash2 speculative decoding). Those speed things up on the Mac alone too; the numbers below keep the two apart. Tested on a MacBook Pro M4 Pro (24 GB) with iPhone 17 Pro Max (A19 Pro) and iPhone 16 Pro Max (A18 Pro) phones.…

StayLameBro/backburner on GitHub A short preview, not the full document.

Read the full README ↗ · Preview checked 2026-10-03T13:16:37.309Z

Inside the original README — Document outline
  1. Backburner
  2. Results
  3. Reading (prefill): Mac alone vs Mac + iPhone, same build
  4. Writing (decode)
  5. Context you can hold
  6. How it works
  7. Who does what
  8. The pieces
  9. Limits
  10. Setup
  11. Reproducing the numbers
  12. Post your results

Headings are captured from the source. Links open the publisher’s document, not a locally hosted copy.

Links from the README

References supplied by the publisher, not independently verified endorsements. Check the destination before downloading files or entering credentials.

Documentation belongs to its respective authors. Reported project/model license: MIT. A listing is not a grant of reuse or training rights. Confirm the document’s own terms at the source.

What this repository does

Your iPhone helps your Mac run a 27B model: faster prompt reading and more context over a USB-C cable

Repository facts

Owner
StayLameBro
Primary language
Objective-C++
Stars
158
Forks
13
Open issues + pull requests
5
License
MIT
Archived
No
Default branch
main
Created
2026-10-01T20:59:56.000Z
Last push
2026-10-03T02:44:10.000Z

Topics and intended use

Owner-supplied topics: apple-silicon, ios, iphone, llama-cpp, llm-inference, local-llm, macos, metal, qwen, sme2, speculative-decoding

Review the README for scope, installation, examples and limitations. We do not run repository code or certify it.

Evaluate before installing

Review licensing and dependencies, inspect recent commits and unresolved issues, and test in an isolated environment before production use. Stars and forks alone cannot answer these questions.

README and project files

Issues and maintenance discussion

Releases and changelog

Source and freshness

Source: GitHub. Metadata observed 2026-10-03T12:05:37.640Z. Daily imports are snapshots, not real-time monitoring.

Popularity and source listings do not establish security, suitability, licensing rights or benchmark performance.

Open original source

Related repositories