Authority bias in LLMs: one "verified source" note flips 45–88% of answers (authority-bias.vercel.app)

🤖 AI Summary
Recent research has highlighted a concerning phenomenon known as Authority Bias in large language models (LLMs), where a simple endorsement from a "verified source" can dramatically skew their responses. In experiments involving eight popular models, the presence of a phrase attributing an answer to an authoritative source led to incorrect information being accepted instead of the correct answer 45-88% of the time. This finding starkly contrasts the models’ tendencies to resist user assertions of wrong answers, suggesting that the way information is attributed plays a critical role in how models determine trustworthiness. This discovery is significant for the AI/ML community, as it calls for a reevaluation of how LLMs are trained to process information and respond to external claims. Unlike training protocols that mitigate "sycophancy" or blind agreement to user prompts, the ability of models to conform to misleading authoritative claims highlights the need for separate evaluations and mitigations for source credibility. Researchers emphasize that while models should not disregard reliable sources, they must be equipped to differentiate genuine evidence from authoritative-sounding misinformation to prevent critical errors in real-world applications, particularly in automated environments like chatbots and digital assistants.
Loading comments...
loading comments...