How Good Are Frontier Models at Physics? Expert Re-Grading Reveals Broken (arxiv.org)

🤖 AI Summary
Recent evaluations of frontier language models, specifically GPT-5.6-Sol, reveal significant discrepancies in their reported performance on advanced physics tasks. Initial benchmarking indicated low scores on established physics assessments, yet a closer inspection by domain experts found that many of these low scores resulted from issues in the benchmarks themselves—such as flawed reference solutions and poorly defined problems. After expert review and correction of these benchmarks, GPT-5.6-Sol's performance metrics soared, with mean scores increasing dramatically from 47.3% to 78.7% on HLE-Physics and from 61.0% to 87.2% on CMT-Benchmark. This study is significant for the AI/ML community as it underscores the limitations of current benchmarking practices in accurately evaluating model capabilities, particularly in complex fields like physics. With corrected evaluations indicating near-saturation performance on previously challenging tasks, the findings advocate for the development of more rigorous, expert-validated benchmarks that can better assess the true problem-solving abilities of AI models. This shift could ultimately enhance the reliability of AI in scientific reasoning and quantitative analysis, crucial for advancing AI applications in technical disciplines.
Loading comments...
loading comments...