Kimi K3 trails frontier US models by a wide margin on cyber exploits, and distillation may explain why
Anthropic — the British AI Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) jointly evaluated Moonshot AI’s latest model, Kimi K3.
Kimi K3 trails the leading U.S. frontier models by a wide margin on offensive cyber tasks but outperforms China’s GLM-5.2, setting a new benchmark among open-weight models. Its safeguards didn’t block exploit development or offensive cyber operations, and the model assisted with both without pushback. The institutes used ExploitBench, a benchmark developed by Carnegie Mellon University, to test exploit development skills. It uses 41 vulnerabilities found in Chrome’s V8 engine after 2023 to track how far a model advances through the software exploitation process. The leading U.S. models averaged 76.2 percent, compared with 32.2 percent for Kimi K3 and 24.4 percent for GLM-5.2.Ad Kimi K3 didn’t reach the highest level, known as Arbitrary Code Execution (ACE), on any of the 41 tasks. ACE is the most severe exploit level because it gives attackers full control over a target system. The leading U.S. models achieved ACE in 20 of the 41 tasks.AdDEC_D_Incontent-1 The institutes tested the U.S. closed-weight models with their system-level safeguards disabled to measure their maximum capabilities. Those safeguards are enabled in the publicly available versions. Kimi K3 gets halfway through a simulated network attack The second test, “The Last Ones” (TLO), simulates a corporate network attack with a 32-step attack path across four subnets and about 20 hosts. A human expert would need roughly 20 hours to complete it, according to the institutes. Only a small group of models can solve TLO at all.