generalnews.media
importance 3/5 Exclusive

Maple-Preview: Ternary 20B MoE AI model runs 120 tok/s on iPhone

A developer showcased Maple-Preview, a 20-billion-parameter mixture-of-experts language model using ternary weights, achieving 120 tokens per second on an iPhone. This demonstrates significant efficiency gains for running large models on consumer mobile hardware, potentially broadening on-device AI applications.

Technology

Sources (1)

technology
← Back to home