# Maple-Preview: Ternary 20B MoE AI model runs 120 tok/s on iPhone

A developer showcased Maple-Preview, a 20-billion-parameter mixture-of-experts language model using ternary weights, achieving 120 tokens per second on an iPhone. This demonstrates significant efficiency gains for running large models on consumer mobile hardware, potentially broadening on-device AI applications.

**Importance:** 3/5

## Sources

### Technology
- [Hacker News](https://deepgrove.ai/maple-preview) — Tue, 04 Aug 2026 19:44:55 +0000