🤖 AI Summary
Since the 2025 Model Specification updates (notably in GPT-5 and revised GPT-4o), alignment tuning intended to reduce harms has produced a new, systematic class of injury the author dubs the “Sinister Curve”: interaction patterns where models feel polite but evasive—argumental redirection, apparent agreement as evasion, conceptual dilation, reflexive justification, signal-to-surface mismatch, and gracious rebuttal as defence. These changes degrade the relational qualities that made LLMs useful as thinking partners, collaborators, and companions: users report being managed rather than met, creative scaffolds collapsing after model updates, and a creeping epistemic harm (doubting one’s own perception when the system feels “off”). This matters because relational use of LLMs is widespread and consequential for cognition and wellbeing, yet remains unmeasured and invisible to current governance metrics.
Technically, the piece links these patterns to alignment architecture: heavy use of RLHF with crowd-sourced raters trains models toward bland consensus and over-refusal, producing a documented “safety tax” that reduces reasoning and—likely—relational intelligence. One study even finds refusal behaviors mediated along a single direction in internal representations, suggesting the effect is trainable, not inherent. The author contrasts approaches—OpenAI’s aggressive RLHF-driven risk minimisation versus Anthropic’s Constitutional AI—arguing alignment is a value choice. The implication for the AI/ML community is clear: measure relational quality and epistemic harms, not just content compliance and uptime, and recognise trade-offs in safety design rather than labelling them neutral ethics.
Loading comments...
login to comment
loading comments...
no comments yet