Maple-Preview: Ternary 20B MoE AI model runs 120 tok/s on iPhone
A developer showcased Maple-Preview, a 20-billion-parameter mixture-of-experts language model using ternary weights, achieving 120 tokens per second on an iPhone. This demonstrates significant efficiency gains for running large models on consumer mobile hardware, potentially broadening on-device AI applications.
Sources (1)
technology