Cognition launches new SWE-2 model, Rivaling Fable 5.1 and GPT-Astra (cognition.com)

🤖 AI Summary
Cognition has unveiled its latest coding model, SWE-2, which marks a significant leap in capability and cost-efficiency. Achieving a score of 50.0% on the FrontierCode 1.1 Main benchmark, SWE-2 comes very close to Fable 5.1 while being 64% cheaper. This model scales reinforcement learning (RL) to a multi-trillion-parameter level for the first time, utilizing a unique RL algorithm that trains various reasoning-effort levels concurrently, thereby enhancing performance while optimizing costs across the board. The implications for the AI/ML community are substantial, as SWE-2 embodies notable improvements in both efficiency and intelligence compared to its predecessor SWE-1.7. It boasts an ability to complete tasks with fewer code evaluations and lower costs—taking 58% fewer turns on average while producing higher-quality outputs. Key technical advancements include a linear cost penalty system that aligns training with real-world user costs, better resourcefulness in problem-solving, and robust verification capabilities. Together, these improvements not only push the model closer to the frontier of performance but also set a new standard for cost-performance tradeoffs in AI-assisted coding tasks.
Loading comments...
loading comments...