Simular's Sai tops OSWorld 2.0, beats GPT and Opus at 2/3 the cost (www.simular.ai)

🤖 AI Summary
Simular's newly launched computer agent, Sai, has claimed the top spot in the OSWorld 2.0 benchmark with a remarkable 73% success rate, surpassing OpenAI's GPT-5.6 Sol and Anthropic's Opus 5 at lower operational costs—approximately two-thirds the price of its competitors. OSWorld 2.0, designed to assess complex real-world tasks that typically require skilled human proficiency, includes 108 diverse challenges, reflecting realistic scenarios like managing dynamic data and executing detailed planning. This achievement signifies a pivotal advancement in AI, demonstrating that enhanced efficiency does not have to come at a higher price. Sai integrates innovative neuro-symbolic planning, which allows it to take fewer actions while maintaining high task throughput. This method enables the agent to orchestrate different specialized models efficiently, optimizing memory usage and processing speed. For instance, in practical tests, Sai completed a gaming task with superior efficiency by treating it like a programming challenge, achieving high scores in significantly fewer steps compared to its counterparts. These advancements mark Sai as a game changer in the Computer Use Agent (CUA) domain, focusing on reliable, cost-effective performance that could revolutionize routine digital tasks. Ultimately, Simular aims to streamline everyday digital responsibilities, allowing users to focus on higher-level decision-making and creativity.
Loading comments...
loading comments...