This 27B reasoning model runs on an iPhone and Apple is testing it
Curated by the Inblix editorial team
PrismML, a startup founded by Caltech researchers, just released Bonsai 27B — a 27-billion-parameter model that runs locally on an iPhone. That’s not a typo. We’re talking about a model that typically guzzles 54 GB of storage, compressed down to as little as 3.9 GB, small enough for an iPhone 17 Pro Max. The secret is aggressive quantization that pushes weights down to just 1–2 bits, an approach that CEO Babak Hassibi confirmed Apple and others are now testing. The talks are “very early,” he told CNBC, but “things are progressing nicely.”
The pitch is straightforward: running AI on-device eliminates per-token cloud costs and keeps private data — screen content, documents, tool calls — from ever leaving the phone. For agent-based workflows that might chain hundreds of model calls, those savings compound fast. PrismML sees this as the foundation for always-on assistants and hybrid systems where only the genuinely hard problems get shipped to frontier models. The smaller variant churns out about 11 tokens per second on an iPhone 17 Pro Max, yielding roughly 67,000 tokens on a full charge before throttling kicks in after five minutes.
Two versions are shipping. The quality-focused variant weighs around 5.9 GB for laptops (actual packages run larger depending on the runtime), and the aggressive one hits that 3.9 GB smartphone target. Across 15 benchmarks, the larger version retains 95% of the original Qwen3.6-27B’s performance. Even the smaller one holds onto 90%, with math and coding described as “virtually unaffected.” Image understanding, instruction following, and agent tool use took the biggest hits under the most aggressive compression. But here’s the kicker: a conventionally compressed Qwen3.6-27B at 9.4 GB scores only 72.7 points on PrismML’s benchmark, while Bonsai’s 3.9 GB variant scores 76.1. More than double the compression, better results.
Why Apple cares should be obvious. At WWDC 2026, the company unveiled a revamped Siri built on Google Gemini technology, and its on-device models have lagged competitors in benchmarks. A licensing deal for PrismML’s compression tech could close that gap. The model weights are already available under Apache 2.0, with support for Apple’s MLX framework and NVIDIA GPUs. PrismML has backing from Khosla Ventures, Cerberus, and Google, with Samsung still in the mix. Next up: applying the same compression to Google’s Gemma series. If this works at scale, the line between what runs on your phone and what needs a data center gets a whole lot blurrier.
💡 Key Takeaways
- PrismML compressed a 27B model from 54 GB to 3.9 GB using 1–2 bit quantization, making local AI agents economically viable on smartphones.
- Apple is actively testing the technology, which could help close the performance gap between its on-device models and cloud-based competitors.
- The smaller Bonsai variant outperformed a conventionally compressed 9.4 GB version of the same base model by nearly 5 percentage points on benchmarks.
- An iPhone 17 Pro Max generated about 67,000 tokens on a full charge with the smaller variant, but throttling began after roughly five minutes of sustained use.
Keep reading: See related articles below for more coverage on this topic.
Get smarter about AI
The sharpest AI news, curated daily. Delivered free to your inbox.