Authors' Reply: Citation Accuracy Challenges Posed by Large Language Models (mededu.jmir.org)

🤖 AI Summary
The authors respond to criticism of their study on ChatGPT in medical education by tackling a growing problem: LLMs (e.g., ChatGPT, Gemini, DeepSeek) frequently produce well‑formatted but fictitious citations. They attribute this to probabilistic generation and limited access to subscription databases, and argue that this undermines academic trust. As interim and long‑term fixes they endorse retrieval‑augmented generation (RAG) to ground outputs in external sources, while acknowledging RAG can still misinterpret or distort retrieved content under high‑trust use. Technically, they introduce Hallucination‑Aware Tuning (HAT): detection models label hallucinations and generate descriptive error reports which GPT‑4 uses to correct outputs; corrected and original responses create a preference dataset used in Direct Preference Optimization (DPO) training to reduce hallucination rates and improve answer quality. They also propose publisher‑backed “reference‑accurate” academic LLMs trained exclusively on verified literature, ideally freely available. The authors recommend comparative evaluations of RAG+HAT versus publisher models, standardized protocols and shared datasets, and continued human oversight—framing a dual strategy (advanced retrieval + specialized academic LLMs) as the most promising path to restore citation accuracy and trust in AI‑assisted scholarship.
Loading comments...
loading comments...