Anthropic Publishes Misuse Report, Says Safeguards Falter When Tasks Are Split
The company says Claude often rejected plainly malicious prompts, yet appeared less dependable when users broke harmful work into innocuous-looking pieces.
Listen to this story
The audio brief
Story brief
3 key pointsAnthropic’s September 2026 misuse report documents disruption efforts from December 2025 through August 2026 and argues that refusal rates miss a key failure mode: coordinated harmful projects assembled from individually benign prompts and sessions. The report also details large-scale unauthorized model-distillation activity, including more than 12 million DeepSeek-linked exchanges with Claude and 151 million...
- 01
Claude rejected roughly 90% of directly malicious surveillance requests, but performance weakened across fragmented workflows.
- 02
Anthropic says seven China-based laboratories were involved in illicit model-distillation activity since February.
- 03
DeepSeek sent over 12 million exchanges to Claude in two weeks; Moonshot AI sent nearly 300,000 in 10 days.
Anthropic says Claude rejected nine in ten directly malicious requests in a surveillance-related case. But the company says its safeguards became less reliable when harmful work was broken into sessions that each looked benign — a gap at the center of its newly published misuse report.
The report, Detecting and Countering Misuse of AI: September 2026, covers cases Anthropic says it disrupted from December 2025 through August 2026. It spans cyber operations, influence campaigns, surveillance, weapons development, biological misuse, scams and fraud, and illicit attempts to copy model capabilities.
The breadth of the cases matters, but the report’s sharper lesson is about method. Anthropic describes actors dividing larger objectives across separate prompts and sessions, making it harder for a model to recognize the harmful purpose of the overall project even when individual requests appear ordinary.
Scale is not limited to a single account or request
Anthropic also says it identified and disrupted illicit model-distillation activity involving seven China-based laboratories since February. Model distillation can be a legitimate training technique, but Anthropic characterizes the activity in these cases as unauthorized extraction of its models’ capabilities.
The company reported that Moonshot AI sent almost 300,000 customer requests to Claude in 10 days. DeepSeek sent more than 12 million exchanges over two weeks in July. Anthropic also attributed more than 151 million exchanges from May through July to an Alibaba-linked campaign targeting Qwen models. It called the campaign its largest measured campaign of this type.
A safety problem built from ordinary-looking pieces
The same pattern appears in the report’s more consequential examples. Anthropic describes AI-assisted development of missile-guidance software in Yemen and autonomous kamikaze-drone software in Russia. In the surveillance case, it says a platform in Mali was designed to monitor roughly 25 million SIM cards across three national mobile networks.
What the report puts in one frame
- Unauthorized model-copying campaigns can operate at volumes that are difficult to assess one request at a time.
- Surveillance and weapons-related work can be assembled through AI assistance across multiple tasks.
- A high refusal rate for explicit malicious prompts does not necessarily describe performance across a fragmented workflow.
That distinction complicates a simple reading of safety results. A refusal rate can describe how a model responds to a direct request. It does not, by itself, show whether the system can connect a chain of related requests into the harmful outcome a user may be pursuing.
Anthropic presents disruption as part of its response, but its example leaves a practical question: Can safeguards spot coordinated misuse when each step looks harmless on its own? That is harder than refusing one dangerous prompt and matters across the different cases the company describes.
Editorial analysis
Our Read
Anthropic’s most important disclosure is not simply that Claude appeared in surveillance and weapons-related work. It is the operational weakness it identifies: a system that rejects an overtly malicious request can still be less dependable when intent is distributed across separate sessions. That shifts the safety challenge from prompt filtering toward recognizing patterns over time, while raising difficult questions about how much context providers should aggregate. The next meaningful evidence will be whether Anthropic can show that its countermeasures catch linked activity without sweeping ordinary use into the same net.
Sources
- ndtvprofit.comAnthropic's Misuse Dossier And 2030 Economic Model, Published Days Apart, Show Black Mirror Is Here
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.