MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses (www.lesswrong.com)

🤖 AI Summary
A recent experiment explored the feasibility of using a multi-user dungeon (MUD) environment, named CrucibleBench, for evaluating large language models (LLMs). Researchers tested 13 LLMs across two social objectives within a gamified setting, revealing that rankings of the models were surprisingly sensitive to the individual scoring components. Notably, when the most classifier-dependent metrics were removed, the correlations among model rankings shifted significantly, illustrating a potential bias in evaluations based on LLM classifications. The aggregate κ scores highlighted the challenges, with some metrics showing as low as 0.04 agreement. This study holds substantial implications for the AI/ML community, suggesting that benchmarks utilizing LLM judges should not only report the stability of rankings but also include thorough audits of agreement among judges. The findings challenge the assumption that higher inference costs correlate with better performance, as an unexpectedly costly model ranked below the median. The researchers openly share their methods and data and are seeking feedback to refine their approach, underscoring the collaborative spirit of AI research. The full documentation is available under an open license, inviting further scrutiny and contributions.
Loading comments...
loading comments...