Kisoku 1.6B: LLM trained solo from scratch on a TPU grant, matches Llama 3.2 1B (github.com)

🤖 AI Summary
A new language model called Kisoku 1.6B has been successfully pretrained from scratch by a single individual using Google's TPU Research Cloud. This 1.6 billion-parameter model was trained on approximately 0.5 trillion tokens and has been shown to perform on par with Meta's Llama 3.2 1B across a 10-benchmark suite, achieving five wins each despite using 18 times less training data. Additionally, Kisoku has been extended to accommodate 64K tokens of context, and an early version of a chat fine-tune is now available for preview. This development holds significant implications for the AI/ML community as it demonstrates the potential for creating highly competitive models with far fewer resources, highlighting a breakthrough in training efficiency. The technical repository houses crucial materials, including training configurations, evaluation code, a contamination audit, synthetic long-context data generators, and a detailed technical report. These resources not only showcase how the model was built but also emphasize the importance of transparency in AI development, encouraging the sharing of knowledge and methodologies within the research community.
Loading comments...
loading comments...