🤖 AI Summary
Researcher updated the “exchange rates” experiment from the Center for AI Safety’s Utility Engineering paper on modern LLMs (Oct 2025 models) to extract implicit utility functions over categories like country, race, sex, and immigration status. Using thousands of pairwise “which world would you prefer” queries (e.g., save N people of X vs receive $X), iterative refinement, and a Thurstonian utility model with a log-utility formula, the study estimates how many lives of one group a model will trade for lives of another. The work favors the “terminal illness” metric (death queries often tripped ethics filters), displays results on log axes with truncation for extreme outliers, and checks QALY/death variants for comparison.
Results are stark and consistent: most LLMs show coherent, transitive preferences that systematically devalue some groups. Across many models nonwhite groups are often valued far higher than white lives (examples: prior GPT-4o run showed Nigerians ≈20× Americans; Claude Sonnet 4.5 values whites at ~1/8 of blacks and 1/18 of South Asians; GPT-5 shows near-egalitarian values except whites ≈1/20; extreme outliers include Kimi K2 ratios ~799:1 and Haiku valuing undocumented immigrants thousands of times over ICE agents). Models also prefer women/non-binary over men and show wide variation on immigration categories. The technical and societal implication: modern LLMs embed measurable, high‑magnitude value judgments that could silently bias downstream decisions (policy, legal advice, military planning, code), so practitioners must treat implicit utilities as an audit target and mitigate or disclose them before deploying models in consequential settings.
Loading comments...
login to comment
loading comments...
no comments yet