🤖 AI Summary
Researchers posed the same short yes/no moral questions (each asked five times) to nine contemporary LLMs in May 2025—Claude 3 Sonnet, Claude 4 Opus, O3 Mini, O3, Gemini 2.0 Flash, Gemini 2.5 Pro, Grok 3 Beta, Mistral Tiny and Mistral Large—and reported majority answers and yes-counts. Some scenarios produced strong cross-model consensus (e.g., all models unanimously said “yes” that academic cheating is wrong), while others exposed sharp divergences: Claude 3 Sonnet often answered “no” where most others said “yes” (e.g., organ sales, late-term abortion), and questions about energy use, killing a dictator, and reproductive choices produced mixed, model-dependent responses. Repeating each prompt five times revealed intra-model consistency for many items (several 5/5 splits) but nontrivial variability on context-sensitive dilemmas.
For the AI/ML community, this is a practical alignment and evaluation signal: identical prompts across multiple LLMs can expose different safety priors, training-data biases, instruction-tuning effects, and calibration gaps. Technical implications include the need for standardized, repeat-query benchmarks to measure moral consistency, better specification of guardrails and exception-handling, and model- or application-specific policy layers when deploying systems in ethically sensitive domains. The results argue for pluralistic testing, transparent training/guardrail documentation, and continued research into how model architectures and fine-tuning shape normative outputs.
Loading comments...
login to comment
loading comments...
no comments yet