🤖 AI Summary
A new benchmark called ActiveVision has been introduced to measure the capability of multimodal large language models (MLLMs) in active observation, an essential aspect of human vision where gaze is redirected based on intermediate hypotheses. This benchmark includes 17 tasks divided into three categories that encourage repeated visual perception rather than relying on static descriptions. The results reveal a significant shortfall in the active observational abilities of leading MLLMs: for instance, GPT-5.5 managed to solve only 10.6% of the items, and Claude Fable 5 just 3.5%, compared to an impressive 96.1% average score from human participants.
The significance of ActiveVision lies in its empirical testing of AI models against the cognitive strategies used by humans, highlighting a critical gap in perception-reasoning integration. The performance issues persisted even when models attempted to write their own vision code, underscoring the unreliability of their outputs when confronted with realistic imagery. These findings suggest a pressing need for novel architectures and training methodologies that enhance the integration of perception and reasoning capabilities in MLLMs, pushing forward the frontier of AI development in mimicking human-like visual cognition.
Loading comments...
login to comment
loading comments...
no comments yet