Show HN: Vev – Ask questions about a screenshot, get a probability per answer (github.com)

🤖 AI Summary
A new AI tool called Vev has been introduced, which enables users to pose yes/no, multiple-choice, or graded questions about visual content such as screenshots and JSON records, returning the probability for each possible answer within approximately 40 milliseconds. Vev operates by segmenting an image into vertical slices and processing each slice for decision-making, demonstrated through applications like real-time gaming analysis, where it can identify game states and actions based on visual input. This capability marks a significant evolution in AI systems by integrating image interpretation within decision-making frameworks, akin to Jev's decision models. The Vev models, available in two sizes—vev-4b and vev-9b—require substantial GPU resources (10 GB and 19 GB of memory, respectively) and are optimized for text and image processing tasks. While Vev's initial accuracy on various benchmarks shows promise, particularly in UI state assessments, it performs less robustly than Jev in certain textual scenarios. As a research preview, it provides an opportunity for developers to explore its potential while encouraging feedback for future iterations. Vev's alignment with the Jev API also facilitates straightforward integration for existing users of TypeSafe’s SDK.
Loading comments...
loading comments...