Vosti: Specifying, Implementing, and Verifying Deterministic LLM Inference (arxiv.org)

🤖 AI Summary
Researchers have introduced Vosti, a new inference engine designed to ensure deterministic outputs in large language models (LLMs). This innovation addresses a critical challenge in the AI/ML community: achieving consistent results across different executions despite optimizations such as batch composition and cache management. Traditional systems, including vLLM and SGLang, strive for this goal but often produce varying outputs due to execution differences. Vosti formalizes a specification where identical prompts lead to identical logits, enhancing reliability in LLM applications. The significance of Vosti lies in its formal verification approach, which guarantees deterministic output while maintaining performance similar to existing high-efficiency models. By independently choosing kernels and aligning cached values with specific token prefixes, it effectively preserves output integrity across diverse execution environments. The research employs advanced proof techniques, including a Verus inductive proof and a Triton analyzer, to validate Vosti's operations. This advancement not only strengthens quality assurance in LLM systems but also opens new avenues for deploying AI models in critical applications where consistent output is paramount.
Loading comments...
loading comments...