Sriti Core – Local-first LLM router that cascades Ollama to cloud to frontier (github.com)

🤖 AI Summary
Sriti Core has launched an innovative open-source intelligent routing and cascading engine for Large Language Models (LLMs) that promises to enhance efficiency and reduce cloud costs. By dynamically categorizing requests and checking an encrypted semantic cache, Sriti Core can route tasks through a tiered model hierarchy, starting from local edge models (Tier 3) to cloud-based providers (Tier 2), and finally to frontier models (Tier 1) only when necessary. This approach not only cuts down the reliance on expensive models but also learns the reliability of various models in real-time, thereby ensuring low latency and high output quality. The significance of Sriti Core lies in its ability to manage model selection intelligently while slashing API costs associated with LLMs. Key features include an in-process task classifier for zero-shot categorization, semantic caching, and a dynamic cascade mechanism with quality gates that automatically escalates requests to higher-tier models when lower tiers fail to meet performance metrics. Additional functionalities like encrypted caching, memory retention of previous routing decisions, and a lightweight setup ensure that Sriti Core can adapt seamlessly in diverse environments, providing a robust infrastructure for developers and organizations leveraging LLM technology.
Loading comments...
loading comments...