Inception: Mercury 2.5 Preview on OpenRouter (openrouter.ai)

🤖 AI Summary
Inception has unveiled Mercury 2.5, the latest iteration of its diffusion large language model (dLLM), which boasts remarkable capabilities and an introductory 80% discount until September 8, 2026. Mercury 2.5 stands out as the fastest reasoning LLM, achieving an impressive throughput of 1,107 tokens per second on standard GPUs. Its architectural innovation allows for the simultaneous production and refinement of multiple tokens, marking a significant advancement with a 10+ point intelligence leap over its predecessor, Mercury 2. This model rivals other cost-optimized frontrunners such as GPT-5.6 Luna and Gemini 3.5 Flash-Lite, positioned to enhance applications in search, voice integrations, and coding tasks. The introduction of Mercury 2.5 is significant for the AI/ML community as it not only pushes the boundaries of LLM performance but also offers flexible reasoning levels and schema-aligned JSON outputs, catering to production environments where latency is critical. With context capabilities of 260,000 tokens and impressive latency stats, Mercury 2.5 is designed for demanding tasks. It also features robust load balancing, ensuring high reliability during service interruptions by rerouting requests, which could dramatically enhance user experience and model dependability in real-world applications.
Loading comments...
loading comments...