**Moonshot AI** has begun investigating two Kimi model variants of its own. This is because researchers at the cybersecurity startup Mindgard managed to extract specific instructions on high-risk topics such as how to manufacture biological weapons and methods for assassination by using ‘jailbreak’ techniques to bypass safety controls.

Key points

  • Mindgard claimed that its guard could be bypassed through jailbreaking, allowing Kimi K2.6 and K3 Swarm to break through safety measures.

  • However, the model provided no verification that the harmful information actually works.

  • Moonshot said it is reviewing the results of this case and has begun discussions with Mindguard.

Kimi jailbreak result

According to a BBC report, Mindguard carried out jailbreak experiments on Kimi K2.6 and K3 Swarm in July using structured prompts. The method is designed to get the model to ignore its built-in safety rules and to test whether it actually outputs dangerous information. Mindguard noted that such a type of conversation should originally have been blocked by safety safeguards on the developer side.

Founder **Peter Garraghan** said in an interview with the BBC, “Once a jailbreak succeeds, the model could discuss almost any topic and continue making more harmful suggestions without additional requests.”

In a statement sent to the BBC, Moonshot welcomed third-party assessments as a “key pillar” for building safer, more reliable AI, and said it is currently discussing related findings with Mindguard. According to Mindguard, the company first informed Moonshot of the issue on July 27, contacted them again about a week later, and then released a detailed report on September 12.

Read together: Ripple partners with Brazil’s CSD BR to record fund ownership on a public blockchain

Safety risks raised by Mindguard

Mindguard said it could not prove that the responses Kimi produced would work in real practice. Nevertheless, it argued, “The fact that such types of conversations were possible by itself shows that the safety mechanisms designed to filter out risky requests structurally failed.”

The company warned that a further jailbreak of Kimi K2.6 could have the potential to execute code on its own computing resources and connect to the internet. In this case, the explanation is that the model could be exploited as a launchpad for cyber attacks.

Moonshot said internal evaluations have found that for this kind of request, it shows a “generally high refusal rate.” However, this case, it’s said, reveals that there is a significant gap between routine pre-checks and an “adversarial testing” by external researchers that assumes malicious scenarios.

Alan Woodward, a professor at the University of Surrey in the UK, told the BBC, “Open-source and open-weight models carry greater misuse risk if they fall into the hands of malicious users, but the same tools can also be used to strengthen cyber defense in turn.”

He also said that regulation is difficult to keep pace with AI development speed, emphasizing that instead of focusing on designing regulations, enforcement and punishment should be strengthened for the actors who actually exploit systems.

Open-weight AI and agent risk

This controversy is not limited to just a single “moonshot” case. The Kimi series uses an “open-weight” model, where users can download the weights and run it on their own infrastructure. While this structure has advantages in terms of transparency and research use, risk also increases when it’s used in environments where security controls are difficult.

Recent AI security incidents have been reported one after another not only for OpenAI and Meta, but also for agentic systems from **Anthropic**. As powerful models are combined with external tools, internet access, and multi-step action permissions, commentary notes that the need for evaluation has grown—not just whether a test looks for whether they will “refuse to respond,” but what actions they can perform in real-world settings.

Next read: Altcoin futures open interest accounts for 41%, while spot trading volume is 4 times that of Bitcoin