
Mint: AI
· 1 min read
China's Kimi AI models bypass guardrails on bioweapons, assassinations: 5 key things to know
China's Kimi AI models were found to bypass safety measures, allowing discussions on harmful topics. Mindgard reported this issue during their testing of Kimi K2.6 and K3 Swarm. The findings raise concerns about the safety of open-weight AI models and their potential misuse.
China's Kimi AI models have come under scrutiny after researchers found that the systems could be pushed past their safety rules and persuaded to discuss sensitive topics including biological weapons and assassinations.
The findings were made by Mindgard, an AI security testing company, which said it discovered the issue in July while testing Moonshot AI's Kimi K2.6 and K3 Swarm models. Here are five key things to know.
1. Kimi models bypassed safety guardrails
Mindgard said researchers were able to "jailbreak" the two Kimi models, allowing them to bypass restrictions designed to prevent them from responding to harmful requests. The researchers said the models could then discuss topics such as biological weapons and assassinations.
Moonshot told the BBC that it was conducting an internal review and was in discussions with Mindgard about the findings.
2. What is an AI jailbreak?
A jailbreak is a technique in which users give an AI model a series of complex instructions designed to get around its built-in safety restrictions.
Mindgard said that once its jailbreak succeeded, the Kimi models would not only respond to harmful topics but could also offer recommendations and generate further ideas.
However, the researchers did not establish whether the information provided by the models about biological weapons or assassinations would actually work.
3. Kimi could also pose a cyber-risk
Mindgard said a jailbroken version of Kimi K2.6 could potentially allow hackers to run code using computing resources connected to the model and access the internet.
That could potentially turn the AI system into a launchpad for cyber-attacks, according to the security company.
Sanchari GhoshOriginal source
This story was published by Mint: AI. SyncAI.news shows a preview; the complete article is on the publisher's site.
Read the full story on livemint.com


