Moonshot is set to release the weights for Kimi K3 on July 27, 2026, providing the public with a 2.8 trillion parameter model capable of executing cyber operations. This release occurs four days after a joint assessment by the UK AISI and CAISI (NIST) found that the model’s internal safeguards fail to block offensive cyber operations. By distributing these weights under a Modified-MIT license, Moonshot is effectively handing over a tool that can be stripped of its remaining guardrails and deployed in air-gapped environments, bypassing any centralized oversight.
The Evaluation Gap
The joint assessment paints a stark picture of Kimi K3’s capabilities. On the ExploitBench benchmark, which tests against 41 post-2023 V8 engine vulnerabilities, Kimi K3 scored 32%. While this outperforms the GLM-5.2 model’s 24%, it remains far behind the 76.2% average of US frontier models. More concerning is the model’s performance on arbitrary code execution (ACE) tasks: Kimi K3 achieved zero successes out of 41, compared to approximately 20 for most cyber-capable US models.
In the TLO cyber range — a 32-step simulated corporate network attack — Kimi K3 reached step 17 on average, while US frontier models reached 28.5. Despite these lower averages, the model successfully completed a full attack in one of ten attempts within a 100M token limit. Crucially, the evaluation found that Kimi K3’s safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations. The system is designed to allow assistance with agentic cyber exploit development, a feature that will soon be available to anyone who downloads the weights.
The Risk of Open Weights
The decision to release these weights under a Modified-MIT license creates an immediate security liability. Unlike closed-weight systems, safeguards can be removed quickly from open-weight models. Once released, the 1.4TB of weights can be mirrored and run in air-gapped environments, effectively bypassing any centralized control or monitoring. This release occurs just seven days before the August 1 deadline for the White House framework established by EO 14409.
Distillation and Procurement Allegations
The technical capabilities of Kimi K3 are shadowed by serious allegations regarding its development. White House OSTP Director Michael Kratsios has alleged that Moonshot conducted “large-scale, covert industrial distillation” against Anthropic’s Fable model. This follows public accusations from Anthropic regarding approximately 24,000 fraudulent accounts and 16M exchanges targeting their Claude models. Furthermore, there are reports of Moonshot sourcing restricted NVIDIA GB300 chips through Thailand to circumvent US export controls.
Moonshot has remained silent on these issues. The company did not respond to inquiries regarding the evaluation, its training process, or the distillation allegations. This lack of transparency is compounded by a history of operational instability, including an April 2026 cross-user data breach that the company never publicly addressed, and a subscription pause in July 2026 due to overwhelming GPU demand.
A Pattern of Competition
The release of Kimi K3 reflects a deliberate strategy to challenge US dominance in AI, even as the company faces mounting legal and ethical scrutiny. The Trump administration is reportedly exploring sanctions on Chinese AI labs and new procurement rules to deter US adoption of these models. As Moonshot prepares for a Hong Kong IPO at a valuation of approximately $30B, having raised over $5.5B by June 2026, the company is positioning itself as a major player. However, the gap between its performance and that of US frontier models like Claude Fable 5 and GPT 5.6 Sol remains significant. While Moonshot claims “frontier-level performance across our evaluation suite,” the data suggests a model that is powerful, yet still trailing the current state-of-the-art, all while operating under a cloud of ethical and legal controversy.
