🤖 AI Summary
Researchers released the first large, multi-lab cross-cultural database of taboo words collected bottom‑up from speakers in 17 countries and 13 languages (N = 1,046 in Study 1). Participants freely listed taboo single- and multiword expressions; native researchers translated items to English, annotated them into categories (insult, slur, sexual, scatological, profanities/blasphemies, etc.), and made the full annotated dataset publicly available. In a follow-up rating study (Study 2; N ≈ 455 per measure) each word was rated on six semantic dimensions to characterize what makes words “taboo.” Analyses used mixed-effects models across the 18 lab samples to isolate cross-linguistic and cross-cultural patterns.
Key findings: across all languages, taboo items cluster as extremely low valence, high arousal, and very low written frequency, distinguishing them from neutral or merely emotional vocabulary. Crucially, there is substantial cross‑country variability in perceived tabooness and offensiveness, underscoring that sociocultural context — not just lexical semantics — shapes what counts as taboo. Implications: the dataset is a valuable resource for psycholinguistics, affective computing, and NLP (e.g., better multilingual content moderation, culturally aware sentiment models), but the authors caution about translation/annotation limits and stress the need for community‑specific approaches when modeling offensive language.
Loading comments...
login to comment
loading comments...
no comments yet