🤖 AI Summary
A major international study led by the BBC and coordinated by the European Broadcasting Union tested more than 3,000 news queries across four leading AI assistants (ChatGPT, Copilot, Gemini, Perplexity), evaluated by professional journalists from 22 public service media in 18 countries and 14 languages. Reviewers scored responses on accuracy, sourcing, separation of opinion and fact, and contextualisation. Overall, 45% of answers contained at least one significant issue: 31% had serious sourcing failures (missing, misleading, or incorrect attributions), 20% had major accuracy problems (hallucinated details or outdated information), and Gemini performed worst with 76% of responses flagged—largely due to poor sourcing.
The findings matter because AI assistants are becoming a primary news gateway (7% of online news consumers overall, 15% of under-25s), yet many users assume their outputs are reliable. Systemic, cross-border and multilingual errors risk eroding public trust and could affect democratic participation. The research team released a News Integrity in AI Assistants Toolkit to define what good news responses should look like and to guide fixes, and they’re urging regulators to enforce information-integrity rules and fund ongoing independent monitoring. Key technical takeaways: improving transparent provenance, timestamping, and robust hallucination mitigation are urgent priorities for developers and newsrooms alike.
Loading comments...
login to comment
loading comments...
no comments yet