🤖 AI Summary
The Axera AX8850 NPU accelerator card is now able to run the Qwen3-0.6B large language model (LLM) directly from GGUF format files thanks to a newly developed custom backend called ggml-axcl. This backend, which integrates with llama.cpp, enables models to operate effectively on the Axera AX8850, which boasts a performance rating of 24 TOPS INT8 and supports 8 GB of LPDDR4x RAM. The innovation allows for weights to be streamed directly from the GGUF during load time, drastically improving efficiency by reducing data traffic and bypassing the typical conversion steps. Currently, the model achieves a decoding speed of 1.3 to 2.7 tokens per second in this setup, contrasting sharply with vendor benchmarks of 13.5 to 16.9 tokens per second for similarly equipped models using baked weights and a closed runtime.
This development is significant for the AI/ML community as it opens up new pathways for LLM performance optimization on low-power embedded systems, such as the Raspberry Pi 5. By facilitating hardware-efficient execution and extensive testing phase verifications, the Axera AX8850 demonstrates a viable route for researchers and developers to leverage advanced AI capabilities in constrained environments. The exploration and experimentation with advanced techniques like cross-fragment QKV fusion also underline the potential for future enhancements in model efficiency and speed. The detailed technical setup and testing framework shared in the announcement offer a blueprint for developers interested in similar integrations.
Loading comments...
login to comment
loading comments...
no comments yet