Anthropic Publishes Misuse Report, Says Safeguards Falter When Tasks Are Split
The company says Claude often rejected plainly malicious prompts, yet appeared less dependable when users broke harmful work into innocuous-looking pieces.
Loading page…
The company says Claude often rejected plainly malicious prompts, yet appeared less dependable when users broke harmful work into innocuous-looking pieces.
Listen to this story
Anthropic’s September 2026 misuse report documents disruption efforts from December 2025 through August 2026 and argues that refusal rates miss a key failure mode: coordinated harmful projects assembled from individually benign prompts and sessions. The report also details large-scale unauthorized model-distillation activity, including more than 12 million DeepSeek-linked exchanges with Claude and 151 million exchanges attributed to an Alibaba-linked campaign targeting Qwen.
Claude rejected roughly 90% of directly malicious surveillance requests, but performance weakened across fragmented workflows.
Anthropic says seven China-based laboratories were involved in illicit model-distillation activity since February.
DeepSeek sent over 12 million exchanges to Claude in two weeks; Moonshot AI sent nearly 300,000 in 10 days.
Anthropic says Claude rejected nine in ten directly malicious requests in a surveillance-related case. But the company says its safeguards became less reliable when harmful work was broken into sessions that each looked benign — a gap at the center of its newly published misuse report.
The report, Detecting and Countering Misuse of AI: September 2026, covers cases Anthropic says it disrupted from December 2025 through August 2026. It spans cyber operations, influence campaigns, surveillance, weapons development, biological misuse, scams and fraud, and illicit attempts to copy model capabilities.
The breadth of the cases matters, but the report’s sharper lesson is about method. Anthropic describes actors dividing larger objectives across separate prompts and sessions, making it harder for a model to recognize the harmful purpose of the overall project even when individual requests appear ordinary.
Anthropic also says it identified and disrupted illicit model-distillation activity involving seven China-based laboratories since February. Model distillation can be a legitimate training technique, but Anthropic characterizes the activity in these cases as unauthorized extraction of its models’ capabilities.
The company reported that Moonshot AI sent almost 300,000 customer requests to Claude in 10 days. DeepSeek sent more than 12 million exchanges over two weeks in July. Anthropic also attributed more than 151 million exchanges from May through July to an Alibaba-linked campaign targeting Qwen models. It called the campaign its largest measured campaign of this type.
The same pattern appears in the report’s more consequential examples. Anthropic describes AI-assisted development of missile-guidance software in Yemen and autonomous kamikaze-drone software in Russia. In the surveillance case, it says a platform in Mali was designed to monitor roughly 25 million SIM cards across three national mobile networks.
That distinction complicates a simple reading of safety results. A refusal rate can describe how a model responds to a direct request. It does not, by itself, show whether the system can connect a chain of related requests into the harmful outcome a user may be pursuing.
Anthropic presents disruption as part of its response, but its example leaves a practical question: Can safeguards spot coordinated misuse when each step looks harmless on its own? That is harder than refusing one dangerous prompt and matters across the different cases the company describes.
Editorial analysis
Anthropic’s most important disclosure is not simply that Claude appeared in surveillance and weapons-related work. It is the operational weakness it identifies: a system that rejects an overtly malicious request can still be less dependable when intent is distributed across separate sessions. That shifts the safety challenge from prompt filtering toward recognizing patterns over time, while raising difficult questions about how much context providers should aggregate. The next meaningful evidence will be whether Anthropic can show that its countermeasures catch linked activity without sweeping ordinary use into the same net.
Loading discussion...
Join the conversation
Share the safeguard you think would matter most.
Be the first to share a perspective or an experience.
Reader comments
Newest comments first. Replies stay oldest first.