Show HN: Peekaboolean – image and jev-like typed questions in typed anwers out (github.com)

🤖 AI Summary
The launch of Peekaboolean, a lightweight vision-language model, marks a significant advancement in AI's ability to answer specific queries about images using boolean responses and choice-based scoring. Designed for rapid operation right on laptops, Peekaboolean accepts an image and context, answering various question types like identifying the image kind, scoring sharpness, and confirming the presence of people—all without generating text. This model leverages the SmolVLM-500M-Instruct architecture, with a frozen vision tower and a calibrated language layer, allowing for accurate responses based on predefined criteria. This model's efficiency and versatility are underscored by its low latency—around 400 ms for six questions on an M1 Pro chip—and its ability to handle the training data from both a robust local teacher model (Qwen3-VL-30B-A3B) and public visual question-answering datasets. The independent scoring system ensures that the model evaluates each question autonomously, enhancing reliability across various tasks without cross-question dependencies. With its Github release, Peekaboolean offers a promising tool for researchers and developers looking to integrate advanced image analysis into their applications while pushing the boundaries of AI interpretability in image processing.
Loading comments...
loading comments...