Distillation, Guardrails and the Race to Exploit: What the Anthropic-Alibaba Case Reveals
On 24 June 2026, a letter sent by Anthropic to the US Senate and the White House became public. Its contents go beyond a mere commercial dispute: they highlight a cyber risk that directly concerns defenders, well beyond the AI sector alone. In it, Anthropic accuses the Chinese giant Alibaba of carrying out what it describes as the largest distillation attack it has ever suffered.
The figures convey the scale of the operation. According to the letter, dated 10 June and addressed to senators Tim Scott and Elizabeth Warren, operators affiliated with Alibaba and its AI lab Qwen allegedly used nearly 25,000 fraudulent accounts to conduct more than 28.8 million exchanges with Claude, between 22 April and 5 June 2026. The campaign reportedly targeted the model's most valuable capabilities: software engineering and agentic reasoning, the very core of what gives Anthropic's frontier models their value. A methodological caveat is in order from the outset: these are, at this stage, one-sided allegations. Alibaba has not responded to press inquiries, separately disputes its inclusion on the Pentagon's blacklist, and the phrase "affiliated operators" is no proof of direct involvement organized by the company.
Distillation, explained simply
Unlike a conventional theft, distillation does not consist of stealing a model's source code or weights. The technique is subtler: it amounts to querying a powerful model en masse, collecting its responses, then using that corpus to train a weaker, cheaper model that imitates its capabilities. The attacker breaks no lock; they watch the master at work millions of times until they reproduce its know-how. That is what makes the practice hard to prevent and costly to detect, and it is why the White House had classified it, as early as April 2026, as a national security issue.
Schematic: the principle of model distillation applied to the Anthropic-Alibaba case.
This is not, moreover, an isolated case. In February 2026, Anthropic had already flagged similar campaigns linked to three other Chinese labs, DeepSeek, Moonshot AI and MiniMax, for a total of more than 16 million exchanges. On its own, the operation now attributed to the Alibaba ecosystem would thus far exceed the three previous ones combined, marking a clear escalation in scale.
The real danger for cybersecurity: a model without guardrails
Beyond intellectual property, the aspect that should hold the attention of any security leader lies elsewhere, and it is explicitly flagged by Anthropic: a model reconstituted through distillation inherits the original's capabilities, but not its guardrails.
To gauge what is lost in the operation, it helps to recall how these guardrails hold up under normal conditions. In a model like Claude, safety is not a filter bolted on at the end: it is forged during training. A long alignment effort, made of human feedback and behavioral principles built into the model, teaches it to recognize a dangerous request and refuse it, whether it is help with making weapons, producing malicious code or intrusion instructions. Around the model sit further controls: analysis of incoming and outgoing requests, terms of use, abuse detection and the closure of accounts that breach them. It is this last link that made it possible to spot the fraudulent accounts behind the campaign. This whole edifice represents a major share of the cost of developing a frontier model.
The problem is that these protections do not transfer automatically during distillation. By training a clone model on the target model's raw outputs alone, one recovers its technical competence while leaving aside its refusals. The theoretical result is a system as capable as the original, but willing to respond to requests the original would categorically refuse.
These guardrails are neither a configuration file one can copy nor an instruction written into a system prompt: they are diffuse across the model's weights, learned during alignment. That is why distillation, which observes only the responses, cannot carry them over.
This is where the subject directly meets defensive cybersecurity. The agentic reasoning and software engineering targeted by the campaign are exactly the skills that allow a model to discover and exploit vulnerabilities at scale. A frontier model endowed with these capabilities, but stripped of its guardrails, would become a highly effective offensive tool in the wrong hands.
The missing link: why attackers will now move even faster
This episode is part of a dynamic our watch has been documenting for months. Over the past month, among all the major cyber signals captured by our system, 302 were directly tied to AI: model security is no longer a forward-looking topic, it is now a continuous news stream. On the defensive side, frontier models have demonstrated a spectacular ability to find flaws: as part of its research program, Anthropic and its partners identified more than 10,000 critical or high vulnerabilities in under two months. But the same finding revealed the real bottleneck: more than 99% of those flaws had not yet been patched, for lack of human capacity to absorb the volume, a blind spot we documented in our article Minor Vulnerabilities: The End of Tacit Risk Acceptance in the Age of AI.
This industrialization of exploitation is anything but abstract, and our watch measures its daily reality. Our mirror of CISA's KEV catalogue currently tracks 1,629 vulnerabilities confirmed as actively exploited in the wild; among them, 327 are already weaponized in ransomware campaigns. These are all flaws for which an exploit exists, circulates and merely awaits chaining, exactly the playground where an offensive model without guardrails would excel.
| Indicator | Value |
|---|---|
| Major cyber signals | ~25,000 |
| of which AI-related security events | 302 |
| Actively exploited vulnerabilities tracked (KEV mirror) | 1,629 |
| of which already weaponized in ransomware | 327 |
Distillation adds a worrying dimension to this picture. If the ability to discover and exploit vulnerabilities at machine speed can be cloned and redistributed into models without guardrails, then the asymmetry between attack and defense worsens further. On the defensive side, one must prioritize, test and deploy patches, a slow and costly process. On the offensive side, an unrestricted distilled model could chain reconnaissance, discovery and exploitation without ethical friction or limitation. This is the concrete translation of a phenomenon defenders already observe: the window between a flaw's disclosure and its exploitation is shrinking to a few days, and the proliferation of offensive capabilities through distillation threatens to shrink it further.
The message for organizations is therefore consistent with the fundamentals: remediation speed is no longer a comfort option. When exploitation capability industrializes and risks spreading beyond any control, the stock of unpatched vulnerabilities an organization lets live becomes an increasingly dangerous debt. Prioritizing by real exploitability, reducing the attack surface and accelerating the patch cycle are the only tenable responses.
Why every company is concerned for its own intellectual property
The lesson does not apply only to AI labs. Any organization that exposes a proprietary model, a specialized assistant or an AI-based service, through an API or a conversational interface, presents the same exposure surface. A product's know-how does not live only in its code or its weights: it shows through in its responses. A competitor or a malicious actor can query that service en masse and reconstitute, at low cost, part of its value, without ever crossing the classic security perimeter. That is what makes the threat hard to grasp: nothing is stolen in the usual sense, no intrusion alert fires, and encryption as well as access control offer no protection here.
For product and security teams, the question shifts accordingly. A model's outputs must be treated as an asset that can be exfiltrated, and therefore monitored. A few measures reduce the risk: limiting and capping the request rate per account, detecting abnormal usage patterns, verifying the identity of high-volume users, explicitly framing distillation in the terms of use and keeping usable records in case of recourse. None of these alone stops a determined campaign, but their absence amounts to leaving a project's value freely copyable by anyone who knows how to ask the right questions, often enough.
A case where the technical and the geopolitical intertwine
It would be naive to read this affair through a purely technical lens. The timeline is dense and revealing. The Pentagon had added Alibaba to its list of Chinese military companies on 8 June, a designation Alibaba is contesting in court. Two days after Anthropic sent its letter, on 12 June, the Department of Commerce imposed export restrictions on Anthropic itself for its most advanced models, Mythos 5 and Fable 5, forcing the company to disable them worldwide, including for its own non-US employees. The shockwave quickly spread beyond the Anthropic case alone: on 25 June, the White House asked OpenAI to restrict the launch of its upcoming model, GPT-5.6, to a small circle of government-approved partners. This was the first time the US government had preemptively gated a model's release, explicitly on the grounds of its cybersecurity capabilities. Anthropic, which is also preparing its IPO, is using this platform to call for tighter export controls and sanctions against distillation practices.
At the level of individual organizations, this wariness is already taking very concrete forms. In late June 2026, France's Directorate General of the Treasury, at Bercy, halted after a few days the trial of an internal assistant built on Qwen, Alibaba's model, the very one at the heart of this affair. Several staff had flagged "oriented or biased answers on China-related topics", and the administration fell back on a French model, citing the model's nationality and the sensitivity of the data involved. The episode highlights the other side of the problem: if distillation can strip from a model what was built into it, namely the guardrails, training can just as easily etch in what one did not choose, that is, its orientations. In both cases these properties are diffuse across the weights and invisible from the outside, which makes vetting the models one deploys a security matter in its own right, not a mere choice of tool.
For a professional reader, the takeaway is not to settle this conflict, but to retain its structural lesson. The security of AI models has become a strategic asset disputed at the highest state level, and the offensive capabilities they hold can spread through unexpected channels. Whether or not the accusation against Alibaba is confirmed, the mechanism it describes is real and reproducible. Adversarial distillation is now part of the threat landscape, just as model theft already featured among the risks catalogued by AI security frameworks.
Mythos, Fable, GLM 5.2: frontier capability leaks faster than it can be locked down
Update, 4 July 2026. The affair only makes full sense in light of what gives the models at its center their value, and their danger. Mythos 5 and Fable 5, the two systems Anthropic had to disable worldwide on 12 June, are not generic assistants: they are its frontier models, the ones whose agentic reasoning and software engineering reach the level that worries regulators. Mythos is precisely the model that, under the Glasswing program, surfaced more than 10,000 critical or high vulnerabilities in under two months. Therein lies the whole stake: the very capability that makes Mythos an unmatched defensive aid would make it, stripped of its guardrails, a first-rate offensive weapon. By classifying these models as export-controlled assets, the state acknowledges that a frontier model's cyber power has become a national security matter, on a par with a dual-use technology.
But a lock is only worth as much as the absence of an open door right beside it. On 3 July 2026, the Chinese vendor Z.ai released GLM 5.2, an open-source model its designers present as rivaling Claude Opus 4.8: barely 1% behind on long technical projects, second overall on complex engineering tasks, and ahead of its Western competitors at, of all things, improving small models, that is, the very ground of distillation. Where Mythos and Fable sit behind an API, terms of use and export controls, GLM 5.2 downloads with "no regional limits" and can be modified freely. Yet a model's guardrails, as we have seen, are diffuse across its weights: an open-weight model can be retrained to strip out its refusals, with no need to mount a 25,000-account distillation campaign. The threat described in this article (a system as capable as the original but willing to do anything) then becomes reachable not at the price of a patient theft, but by a simple download.
This is the uncontrolled-drift risk that extends the Anthropic-Alibaba affair. Within the same fortnight, the United States disables Mythos 5 and Fable 5 worldwide and preemptively gates the launch of GPT-5.6, while an open-weight model of comparable capability spreads with no border or condition. Export control has a grip on the closed frontier; it has none on weights already published. For a defender, the consequence is direct: exploitation at machine speed no longer depends on access to a well-guarded proprietary model. It is now a file one downloads, tunes and runs, which only sharpens the urgency, hammered above, of shortening the remediation cycle before the asymmetry widens further.
Sources
- CNBC - Anthropic accuses Alibaba of distillation campaign (28.8M exchanges, 25,000 accounts, 22 April-5 June 2026) - cnbc.com
- Reuters / Bloomberg - 10 June letter to senators Tim Scott and Elizabeth Warren; targeting of software engineering and agentic reasoning
- The Next Web - Pentagon context (8 June list), Alibaba's lawsuit, Anthropic's IPO - thenextweb.com
- Reuters via Global Banking & Finance - 12 June export restrictions on Mythos 5 and Fable 5
- Euronews Next - GLM 5.2, the new Chinese AI model rivaling Anthropic: Z.ai releases GLM 5.2 as open source on 3 July 2026; ~1% behind Claude Opus 4.8 on long technical projects, 2nd on complex engineering, 1M-token context, "no regional limits" - euronews.com
- Axios / CNN - White House asks OpenAI to restrict GPT-5.6's launch to approved partners (25 June 2026) - axios.com
- Le Monde Informatique - Bercy unplugs a Chinese LLM trial: France's Directorate General of the Treasury halts a Qwen-based (Alibaba) assistant trial after "oriented or biased answers on China-related topics" (26 June 2026) - lemondeinformatique.fr
- futunn / cybersecuritynews - recap of the February 2026 campaigns (DeepSeek, Moonshot, MiniMax, ~16M exchanges) and the "allegations" caveat
- White House / OSTP - April 2026 memo classifying distillation as a national security issue
- Anthropic, Project Glasswing (more than 10,000 critical/high vulnerabilities identified in under two months; >99% unpatched): anthropic.com/glasswing
- Primary data: LuksMentis watch (June 2026); ~25,000 major cyber signals, 302 of them directly tied to AI; mirror of CISA's KEV catalogue: 1,629 actively exploited vulnerabilities, 327 of them ransomware-linked.
- LuksMentis, Minor Vulnerabilities: The End of Tacit Risk Acceptance in the Age of AI: luksmentis.com/blog