UK AISI / Caisi Preliminary Assessment of Kimi K3's Cyber Capabilities (www.nist.gov)

🤖 AI Summary
The UK Artificial Intelligence Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) have released a preliminary assessment of Moonshot AI's Kimi K3 model, unveiling its performance regarding cyber capabilities. Released on July 16, 2026, Kimi K3 was evaluated on benchmarks such as ExploitBench, focusing on exploit development tasks. While it outperformed the open-weight model GLM-5.2 with a score of 32%, it struggled significantly in achieving high-severity exploit outcomes, failing to execute arbitrary code execution (ACE) on any of the 41 tested vulnerabilities, whereas leading models averaged 20 successes out of 41. The implications of this evaluation are significant for the AI/ML community, as Kimi K3 demonstrates both advancements and limitations in automated cyberattack capabilities. Although it showed some potential by successfully completing a portion of the “The Last Ones” (TLO) cyber range, it fell short compared to U.S. models, averaging only 17 out of 32 steps completed in simulated attacks. The assessment highlights the need for ongoing improvements in AI systems for cybersecurity, especially as malicious actors may leverage AI tools for exploitation. The benchmarks utilized provide a framework for continuous accountability and measurement in AI cyber capabilities, crucial for developing more resilient systems against emerging threats.
Loading comments...
loading comments...