Cognition's SWE-2 achieves 92.8 on Terminal-Bench 2.1 (tokenstead.ai)

🤖 AI Summary
Cognition has announced that its latest model, SWE-2, has achieved an impressive score of 92.8 on Terminal-Bench 2.1, marking a significant milestone in scaling reinforcement learning (RL) to the multi-trillion-parameter range, featuring 2.8 trillion total parameters with 104 billion active parameters per token. This accomplishment builds on the already RL-focused Kimi K3 base, enhancing its performance benchmarks by 5 to 6 points across various tasks. The model employs MoE inference on advanced NVFP4 and FP8 kernels and includes innovations like quantization-aware training, which contribute to improved efficiency and performance metrics. The significance of SWE-2 lies not only in its record-setting benchmarks but also in its cost-effectiveness, reportedly offering a 64% lower operating cost compared to Claude Fable 5.1 at similar performance levels. Notably, its mean steps per run are significantly lower, achieving faster edit times with 58% fewer runs than its predecessor SWE-1.7. Despite its strong performance, SWE-2 still faces challenges in long-horizon tasks, as indicated by its lower score on Terminal-Bench 4.0. While SWE-2 is currently not available for local downloads, it can be accessed through Cognition’s platforms, emphasizing its proprietary nature and the ongoing trend towards robust cloud-based AI solutions.
Loading comments...
loading comments...