When “no” means “yes”: Why AI chatbots can’t process Persian social etiquette (arstechnica.com)

🤖 AI Summary
Researchers led by Nikta Gohari Sadr released a study and the TAAROFBENCH benchmark showing that mainstream LLMs systematically fail to handle taarof, the Persian ritual of polite refusal and insistence. Tested models—GPT‑4o, Claude 3.5 Haiku, Llama 3, DeepSeek V3 and even Dorna (a Persian‑tuned Llama 3 variant)—only navigated taarof scenarios correctly 34–42% of the time, while native Persian speakers scored 82%. Taarof involves repeated offers and counter‑refusals, deflecting compliments, and ritualized politeness where what is said often differs from what is meant; current models default to Western‑style directness and miss these implicit cues. The gap has concrete implications for deploying AI in multicultural settings: cultural misreads can derail negotiations, harm relationships, and entrench stereotypes. Technically, the results expose limits in both pretraining and typical fine‑tuning approaches—language adaptation alone doesn’t guarantee pragmatic competence. The TAAROFBENCH provides a targeted evaluation to drive improvements in cultural pragmatics, suggesting the need for culturally aware datasets, instruction tuning focused on indirect speech acts, better context modeling of ritualized exchanges, and human‑in‑the‑loop evaluation by native speakers when local norms matter.
Loading comments...
loading comments...