Is Physics Dead: Broken benchmarks and re-evaluating frontier models in physics (jsous.github.io)

🤖 AI Summary
A recent study highlights significant discrepancies in the evaluation of AI models' capabilities in solving physics problems, revealing that common benchmarks may be misleading. The researchers audited various physics questions after frontier AI models, like GPT-5.6 Sol and GPT-6 Astra, received poor scores—32% on CritPt and 47% on Humanity’s Last Exam. Their investigation uncovered that many previous answers rejected by the models were actually correct, stemming from flawed benchmark items and grading processes. After a thorough audit and corrections, AI performance dramatically improved, with some models approaching saturation on even challenging physics problems, suggesting that AI's proficiency in physics may be underestimated. This finding has critical implications for the AI/ML community, indicating that current benchmarks do not accurately reflect the capabilities of frontier models in physics research. The study advocates for the development of new benchmarks that emphasize problem representation, encourage collaboration between humans and AI on open physics questions, and create verification tools tailored to the nuances of physics. As AI continues to advance, the interaction between human intuition and AI's computational strength may become crucial in redefining future physics challenges, suggesting that the field is far from "dead." Instead, it calls for a proactive approach to harness AI's potential in a way that complements human understanding and creativity in physics.
Loading comments...
loading comments...