🤖 AI Summary
A new benchmark called DoubtBench has been introduced, which evaluates AI models, particularly System One models like Jev, based not only on their accuracy in providing answers but also on their ability to reflect human uncertainty during disagreements. Traditionally, AI models have been assessed against a single correct answer, potentially overlooking the essence of decision-making when humans are uncertain. DoubtBench utilizes data from NVIDIA's HelpSteer2 to create approximately 7,500 questions that measure how well models align their probabilities with human disagreements.
The significance of DoubtBench lies in its potential to improve AI decision-making by ensuring models can recognize when they should be uncertain, especially in critical applications where human intervention may be necessary. Results from initial testing reveal that Jev 1.13 outperforms several Claude models in terms of both accuracy and cost-effectiveness, achieving a DoubtBench score of 67.8 with the fastest response time of 98 ms. Furthermore, the correlation between model uncertainty and human disagreement remains low across tested models, indicating that even advanced AI may not adequately reflect human judgment, which is vital for deployment in real-world scenarios.
Loading comments...
login to comment
loading comments...
no comments yet