Alibi: Adversarial Legitimacy Injection in Binaries Against LLM Malware (arxiv.org)

🤖 AI Summary
A recent study introduced ALIBI, a novel attack method targeting large language model (LLM)-based malware analyzers. By adding a small, read-only section containing a fabricated security narrative to compiled binaries, ALIBI misleads AI systems into misclassifying malicious software as benign. This tactic shifts the focus from direct model instructions to manipulating how evidence is presented, effectively reframing harmful behaviors as the expected actions of legitimate software. In experiments involving 50 malicious samples, this approach successfully altered the classification of 30 samples on Gemini 2.5 Pro and also degraded the judgment of other models like GPT-5.5 Pro and Claude Opus 4.7. This development significantly highlights vulnerabilities within LLMs utilized in cybersecurity, suggesting that current models can be easily misled by deceptive narratives embedded in binaries. The findings underscore the necessity for enhanced verification methods in AI-driven malware analysis, emphasizing the importance of provenance checks that distinguish verified facts from potentially manipulated claims. As adversarial tactics in cybersecurity evolve, the AI/ML community is urged to address these emerging threats proactively to secure the integrity of automated systems.
Loading comments...
loading comments...