Routing LLM traffic across inference providers with TCP-style congestion control (getunblocked.com)

🤖 AI Summary
A new adaptive router has been implemented to efficiently route large language model (LLM) traffic across various inference providers based on cost, speed, and reliability, eliminating the need for manual intervention by engineers. As the market for open-weight models grows increasingly competitive, providers now offer the same models at varying prices and performance levels. This router automatically adjusts traffic to the best available provider, optimizing for cost-effectiveness while taking reliability into account, thereby significantly improving performance and reducing operational overhead. The router functions by scoring providers based on cost and speed metrics calculated per task, ensuring that every request is routed through the most efficient provider. It uses a unique scoring equation that prioritizes cost—accounting for 70% of the score—while still factoring in speed, thus limiting the additional costs for faster services. The system is designed to remain stable, even during provider outages or rate-limit errors, by utilizing a feedback mechanism reminiscent of TCP congestion control that adjusts allowed request rates based on real-time performance data. This innovative approach not only enhances performance but also ensures a seamless user experience without frequent manual adjustments, highlighting the potential for more automated systems in managing AI infrastructure.
Loading comments...
loading comments...