Will We Know Artificial General Intelligence When We See It? (spectrum.ieee.org)

🤖 AI Summary
The article argues that the old Turing test no longer serves as a reliable gauge of artificial general intelligence (AGI) and surveys current efforts to build a more meaningful “IQ test” for machines. Because AGI is variably defined—by performance, internal mechanisms, economic impact or even “vibes”—benchmarking matters: it shapes research priorities, policy and societal preparedness. Human-style intelligence mixes fluid (on‑the‑fly problem solving) and crystallized (stored knowledge) components, and many modern large language models (LLMs) excel at the latter after massive pretraining but fail on tasks requiring rapid abstraction, social reasoning or embodied causal understanding. The article warns about false positives (Clever Hans–style shortcuts) and false negatives from mismatched tests. François Chollet’s Abstraction and Reasoning Corpus (ARC) and its new ARC‑AGI‑2 contest are highlighted as a focused attempt to measure fluid intelligence: hundreds of visual grid puzzles require learning abstract rules from a handful of examples and applying them to novel cases. ARC-AGI-2 ups the difficulty and offers $1M in prizes for systems that solve 85% of 120 puzzles within strict compute limits (four GPUs for 12 hours). Humans average ~60%; current top AI ≈16% (though an unreleased OpenAI o3 variant reportedly hit 88% at an estimated ~$20,000 per puzzle). The piece concludes that such benchmarks are valuable directional tools but remain incomplete—especially for social, physical and multi‑modal intelligence—and that ongoing, evolving benchmarks will be essential to track true progress toward AGI.
Loading comments...
loading comments...