When LLM judges agree, should we believe them? (www.amazon.science)

🤖 AI Summary
A recent paper presented at the International Conference on Machine Learning (ICML) introduces a novel approach to evaluating outputs from large language model (LLM) judges, focusing on the issue of correlated outputs that can misrepresent the diversity of opinions. The authors propose a dependence-aware label aggregation method utilizing Ising models to account for pairwise dependencies among judges. This innovative approach aims to distinguish between independent evidence and shared mistakes, enhancing the reliability of LLM evaluations in tasks like relevance classification, toxicity detection, and summarization assessment. Significantly, the study demonstrates that this new aggregation method yields 9-14% accuracy improvements over traditional weighted majority voting systems across multiple tasks. By treating judge panels as networks, the model assesses both individual judge reliability and the relationships between judges, addressing potential biases stemming from shared training data or prompts. This advancement not only improves accuracy but also provides practical guidance for LLM-as-a-judge frameworks, emphasizing the importance of evaluating panel diversity and recognizing the effects of correlated decisions. Overall, the findings challenge assumptions about LLM outputs and highlight the need for sophisticated aggregation techniques to ensure valid results.
Loading comments...
loading comments...