🤖 AI Summary
The Artificial Analysis Intelligence Index has launched version 4.3, introducing significant upgrades including Terminal-Bench v4 and a new benchmark called AutomationBench-AA that enhances agentic workflow assessments. This version continues the transition towards the forthcoming Intelligence Index v5, focusing on real-world problem-solving by incorporating more complex tasks and a private test set. Key advancements include a recalibration of difficulty for coding tasks, raising the evaluation weight for agentic workflows to 45%, and replacing the 𝜏³-Banking benchmark.
For the first time, the index highlights the performance of advanced models such as Claude Fable 5.1 and GPT-6 Astra, both scoring 53 points, while also shedding light on cost-efficiency—GPT-6 Astra proves to be substantially cheaper at $3.26 per task compared to Claude Fable 5.1 at $7.63. Open weights models like GLM-5.3 Flash lead the ranks further demonstrating a diverse landscape of performance and cost, which is crucial for developers and researchers aiming to optimize AI solutions for various applications. The updates signify a pivotal moment in benchmarking, pushing forward the standards for AI capabilities in real-world scenarios.
Loading comments...
login to comment
loading comments...
no comments yet