Running Gemma 4 26B at 5 tokens per second on a 13-year-old CPU
A developer demonstrated running Google's Gemma 4 26B model at just 5 tokens per second on a 13-year-old Xeon processor without a GPU. This showcases the extreme optimization needed to run large language models on legacy hardware, though performance is impractically slow for real-world use. It highlights ongoing efforts to make AI models more accessible on low-resource devices.
Sources (1)
technology