Ask a model if code is malicious and it reaches for its morals (www.manifold.security)

🤖 AI Summary
Recent research has explored how modern language models, particularly those based on a mixture of experts architecture, determine if code is malicious by routing queries through specific internal networks. When asked about malicious code, the models exhibit significant overlaps with their routing for moral questions, indicating they engage a moral assessment framework while analyzing potential malice. This suggests that language models do not merely evaluate technical parameters but also grapple with ethical implications, effectively weaving morality into their responses. This discovery is significant for the AI/ML community as it highlights the complex interplay between technical evaluation and moral reasoning in AI. By analyzing the routing patterns of models like OLMoE and DeepSeek, researchers found that the expert networks activated for queries about malice were closer to those engaged for moral questions than for purely technical ones, such as legality or vulnerability. Exploring these pathways provides insights into how AI systems internalize concepts of morality and intent, raising important questions about accountability and transparency in AI-driven decision-making. Understanding these mechanisms could inform the development of more nuanced AI systems capable of ethical reasoning.
Loading comments...
loading comments...