I'm (mostly) picking models on speed now, not intelligence (martinalderson.com)

🤖 AI Summary
In a notable shift within the AI/ML community, many professionals are now prioritizing the speed of models over their intelligence for daily tasks. The author reflects on the performance of models around the Opus 4.6 level, which are deemed "smart enough" for various applications, from coding to data analysis. While initial excitement around newer models like Fable waned due to their slower performance, the trend indicates that speed may soon outweigh intelligence as a primary criterion. The observation that models processing at 100-200 tokens per second (tok/s) feel almost instant suggests a growing demand for quicker response times in practical use cases. This speed-centric focus has significant implications for the competitive landscape of AI models. As developers prioritize and innovate around speed, the gap between capabilities of different models narrows, especially among open-weight models. The price for advanced models is decreasing as competition rises, exemplified by OpenAI's 80% price cut for their Luna variant ahead of the DeepSeek V4 Flash GA release. With upcoming advancements in hardware expected to improve output speeds significantly, the landscape may soon see models capable of 500 tok/s and beyond by 2027, potentially reshaping the benchmarks for both speed and intelligence in AI applications.
Loading comments...
loading comments...