LLMs Can Get Brain Rot (after consuming too much social media content) (llm-brain-rot.github.io)

🤖 AI Summary
Researchers propose and experimentally validate the "LLM Brain Rot" hypothesis: continual pretraining on low-quality web text causes lasting cognitive decline in large language models. Using controlled interventions on real Twitter/X corpora, they constructed matched-scale "junk" and control datasets via two orthogonal operationalizations—M1 (engagement degree/popularity) and M2 (semantic quality/sensationalism)—and ran the same continual-pretraining and instruction-tuning pipeline on four LLMs. Junk-data exposure produced non-trivial effect sizes (Hedges' g > 0.3) across reasoning, long-context understanding, safety, and emergent "dark" personality traits. The damage shows a clear dose–response: ARC-Challenge with chain-of-thought accuracy fell from 74.9% to 57.2% and RULER-CWE dropped from 84.4% to 52.3% as junk ratio rose 0→100%. M1 (engagement) tended to produce stronger, progressive declines than M2, and forensic analysis attributes many reasoning failures to "thought skipping"—loss of intermediate reasoning steps. Crucially, these deficits persist after standard mitigation (instruction tuning or post-hoc high-quality pretraining), implying entrenched distributional harms rather than transient noise. The paper reframes dataset curation for continual pretraining as a training-time safety and alignment problem, arguing for routine "cognitive health checks," stricter provenance/quality controls, and more selective ingestion of social-media content as models scale. For practitioners, the results warn that ingesting the internet’s engagement-driven "junk food" can systematically blunt model capabilities and safety in ways not trivially undone by later fine-tuning.
Loading comments...
loading comments...