Run Qwen3.8 27B locally: real numbers from my Mac Studio (terminalbytes.com)

🤖 AI Summary
Qwen3.8 27B, a new local model, has been running effectively on a Mac Studio, providing functionalities like summarizing RSS feeds and organizing PDF scans. This model, notable for its multimodal capabilities and a large 262,144-token context window, generated excitement in the AI community thanks to its utility beyond mere experimentation. Benchmark testing revealed that while Qwen3.8 runs at about 14 tokens per second, it utilizes significantly fewer tokens per answer than its predecessor, Qwen3.6, making response times comparable despite the lower output speed. The implications for the AI/ML community are substantial as Qwen3.8 demonstrates that local models can handle tasks traditionally reserved for cloud solutions with impressive efficiency. Its hybrid attention architecture and robust performance metrics (61.7 on SWE-bench Pro and 89.2 on GPQA Diamond) position it as a powerful tool for developers looking to integrate AI into their workflows without relying on external APIs. Moreover, the detailed benchmarking highlights the necessary hardware requirements, suggesting that local deployment is increasingly feasible for users equipped with mid-range machines, marking a significant shift toward self-hosted AI solutions.
Loading comments...
loading comments...