I got 2.2x more tokens per second from llama.cpp on Intel Arc (grigio.org)

🤖 AI Summary
A recent benchmarking of the llama.cpp framework on a Xiaomi Book Pro 14 revealed a significant boost in performance when utilizing the Intel Arc B390 GPU. By adjusting the model's CPU_MOE (Mixture of Experts) parameter from the default setting, users experienced up to 2.2 times faster token generation and 2.3 times quicker prompt processing, highlighting the importance of optimizing GPU resources for large language models. With a configuration that allowed full GPU offload, the system managed to process a 400-token completion at 33-36 tokens per second, a dramatic increase from the previous 15-17 t/s under CPU_MOE constraints. This finding is crucial for the AI/ML community, as it underscores the viability of leveraging integrated graphics with unified memory for intensive tasks traditionally relegated to discrete GPUs. The research also identified optimal settings for speculative decoding and batch sizes, showing that a sweet spot of two drafts enhances throughput without incurring memory costs. Overall, these insights advocate for re-evaluating model configurations and memory management strategies to maximize performance, especially as language models evolve in complexity and size.
Loading comments...
loading comments...