Sub-1-Bit LLM Compression via Latent Factorization (github.com)

🤖 AI Summary
The development of LittleBit and its successor LittleBit-2 introduces a groundbreaking approach to compressing large language models (LLMs) into the sub-1-bit range, allowing for data storage efficiency without sacrificing performance. By factorizing dense weight matrices into low-rank latent factors and using lightweight learned scales to restore magnitude information, LittleBit achieves an astonishing compression level of 0.1 bits per weight. LittleBit-2 enhances this technique by correcting latent geometry misalignment during initialization through a method known as Joint Iterative Quantization (Joint-ITQ). Importantly, this enhancement maintains the original model architecture during inference, ensuring no additional overhead in model deployment. The significance of these advancements lies in their potential for enabling LLM deployment in resource-constrained environments, thereby democratizing access to high-performing AI tools. The approach is also compatible with Quantization-Aware Training (QAT), which enhances the robustness of these tiny models under training conditions. With support for a variety of popular architectures including Llama and Qwen, and a user-friendly implementation, LittleBit-2 stands to impact both research and practical applications significantly in AI, making high-performance models more accessible.
Loading comments...
loading comments...