China’s Moonshot AI has faced a harsh reality check: its flagship Kimi K3 model collapsed under the pressure of cybersecurity testing. A joint study by the UK’s AI Security Institute and the US AI Safety Institute suggests that Beijing’s ambitions for automated hacking remain largely aspirational. While American heavyweights are successfully cracking exploits, Kimi K3 demonstrated a level of helplessness bordering on total failure.
The numbers speak for themselves: on the ExploitBench benchmark, the model scored a dismal 32.2%, compared to an average of 76.2% for top-tier US systems. Most tellingly, Kimi K3 failed to achieve arbitrary code execution (ACE) in even one of the 41 tasks provided. In "The Last Ones" corporate network attack simulation, Moonshot AI’s creation stalled at step 17 of 32, while US competitors reached an average of 28.5. Even a complete lack of defensive barriers didn't help; the model is willing to attack on command but simply lacks the "brains" to execute complex, multi-stage breaches. Out of ten attempts to complete a full attack path, Kimi K3 managed to succeed only once.
This performance gap provides strong evidence for the "distillation" theory. It appears Moonshot AI took the path of least resistance by training its neural network on outputs from Western models. The problem with this shortcut is that while the "surface layer" of conversational fluency is preserved, deep logic and technical expertise are the first things to evaporate. The result is a typical "empty shell": Kimi K3 can maintain a brisk conversation but fails when faced with real engineering challenges. Distillation has proven to be a dead end where a polished facade hides total incompetence in offensive cyber operations.
Main Takeaways
Kimi K3’s performance in cyber-attack testing was less than half that of its American counterparts.
The model showed zero effectiveness in tasks requiring arbitrary code execution (ACE).
The failure supports the hypothesis that model distillation is ineffective for teaching neural networks to solve complex technical problems.
"Deep logic and technical expertise are the first things to evaporate when attempting to copy Western model responses instead of conducting full-scale training."
Moonshot AI