PerceptionBench: Evaluating Atomic Visual Perception in Multimodal LLMs (www.kimi.com)

🤖 AI Summary
PerceptionBench has been launched as a groundbreaking benchmark designed to evaluate the atomic visual perception capabilities of multimodal large language models (MLLMs). Developed by the Kimi Team, it isolates visual perception into ten atomic categories, derived from observed failures across 42 existing benchmarks. This innovative approach not only highlights specific shortcomings in model perception—with none of the 16 tested MLLMs achieving even 60% accuracy—but also emphasizes the prevalence of perception-related hallucinations as a significant area of weakness. The benchmark consists of 3,000 verified questions targeting individual perceptual capabilities, ensuring that the evaluation relies purely on visual perception rather than reasoning or external knowledge. By employing a failure-driven taxonomy, PerceptionBench offers a cohesive assessment of visual perception errors that are often overlooked in traditional benchmarks. This capability-centric evaluation promises to drive advancements in multimodal AI, ultimately leading to models that can perceive accurately and consistently. The open-sourcing of the PerceptionBench dataset and evaluation tools invites the AI/ML community to collaboratively address these critical gaps in visual perception.
Loading comments...
loading comments...