Valen: A multimodal decision model inspired by Jev (github.com)

🤖 AI Summary
A new multimodal decision model named Valen has been announced, designed to integrate visual perception into decision-making processes. Inspired by the Jev model, Valen processes text, images, and video inputs to assess candidates based on task instructions, outputting decision probabilities without generating answer tokens. Utilizing a Qwen3.5-0.8B or 2B backbone, Valen provides a structured decision interface that significantly enhances the efficiency and accuracy of system responses. The significance of Valen lies in its ability to perform complex decision-making in a fraction of the time compared to existing models, achieving rapid decision-making with low latency. For instance, the Valen-Preview-0923 version solved a puzzle in just 9 decisions over 1.13 seconds, whereas the Qwen3.8 struggled with significantly longer computation times. Moreover, Valen demonstrates resilience in decision-making confidence under varying image clarity, maintaining a 91.6% confidence on clear images, but dropping to 19.2% with heavy Gaussian blur. This ability to adapt decisions based on visual input clarity presents important implications for applications in real-time AI systems and interactive environments, encouraging further exploration and contributions to the model’s development.
Loading comments...
loading comments...