Everyone prefers human writers, even AI (arxiv.org)

🤖 AI Summary
Researchers ran controlled experiments using Raymond Queneau’s Exercises in Style to measure how humans and AI judge literary style when passages are either original or GPT-4–generated and presented under three labeling conditions (blind, correctly labeled, counterfactually labeled). Study 1 tested 556 human participants and 13 AI models; Study 2 extended this to a 14×14 matrix of AI creators and evaluators. Both studies found a consistent pro-human attribution bias: humans favored human-written text by +13.7 percentage points (Cohen’s h = 0.28, 95% CI 0.21–0.34), while AI evaluators showed a much larger +34.3 pp bias (h = 0.70, 95% CI 0.65–0.76), a 2.5× stronger effect (P < 0.001). Study 2 showed the effect generalizes across architectures (+25.8 pp, 95% CI 24.1–27.6%). Crucially, labeling alone flipped evaluative criteria—identical features were judged oppositely depending only on whether they were labeled “human” or “AI.” This work is significant because it demonstrates that AI systems not only mimic human cultural prejudice against machine-generated creativity but amplify it, with implications for benchmark design, content moderation, and human-AI collaboration. Attribution labels can systematically skew evaluations and create feedback loops if AI trained on biased judgments reinforces those biases. The findings call for calibration of AI evaluators, careful treatment of authorship labels in datasets and benchmarks, and research into debiasing strategies to prevent automated amplification of cultural preferences.
Loading comments...
loading comments...