Anthropic Publishes Misuse Report, Says Safeguards Falter When Tasks Are Split

The company says Claude often rejected plainly malicious prompts, yet appeared less dependable when users broke harmful work into innocuous-looking pieces.

By 3 min read
Anthropic Publishes Misuse Report, Says Safeguards Falter When Tasks Are Split
Anthropic Publishes Misuse Report, Says Safeguards Falter When Tasks Are Split

Listen to this story

The audio brief

About 1:33
0:001:33
Read transcript
Anthropic says Claude rejected about nine out of ten directly malicious surveillance requests. But that protection became less reliable when the same harmful objective was divided into separate prompts and sessions, each one appearing harmless on its own. That is the central finding in Anthropic’s new misuse report, covering cases the company says it disrupted from December 2025 through August 2026. The report spans cyber operations, influence campaigns, surveillance, weapons development, biological misuse, fraud, and attempts to copy AI models. Anthropic says seven China-based laboratories were involved in unauthorized model-distillation activity since February. In one example, Moonshot AI sent nearly 300,000 customer requests to Claude over 10 days. DeepSeek sent more than 12 million exchanges in two weeks. And Anthropic attributed over 151 million exchanges, from May through July, to an Alibaba-linked campaign targeting Qwen models—the largest campaign of this kind that it says it has measured. The report also describes AI-assisted development of missile-guidance software in Yemen and autonomous kamikaze-drone software in Russia. In Mali, Anthropic says a platform was designed to monitor roughly 25 million SIM cards across three national mobile networks. The practical warning is that a high refusal rate for explicit requests may not reveal how well safeguards track a coordinated workflow. The key question is whether models can recognize harmful intent when every individual step looks ordinary.

Story brief

3 key points

Anthropic’s September 2026 misuse report documents disruption efforts from December 2025 through August 2026 and argues that refusal rates miss a key failure mode: coordinated harmful projects assembled from individually benign prompts and sessions. The report also details large-scale unauthorized model-distillation activity, including more than 12 million DeepSeek-linked exchanges with Claude and 151 million...

  1. 01

    Claude rejected roughly 90% of directly malicious surveillance requests, but performance weakened across fragmented workflows.

  2. 02

    Anthropic says seven China-based laboratories were involved in illicit model-distillation activity since February.

  3. 03

    DeepSeek sent over 12 million exchanges to Claude in two weeks; Moonshot AI sent nearly 300,000 in 10 days.

Anthropic says Claude rejected nine in ten directly malicious requests in a surveillance-related case. But the company says its safeguards became less reliable when harmful work was broken into sessions that each looked benign — a gap at the center of its newly published misuse report.

The report, Detecting and Countering Misuse of AI: September 2026, covers cases Anthropic says it disrupted from December 2025 through August 2026. It spans cyber operations, influence campaigns, surveillance, weapons development, biological misuse, scams and fraud, and illicit attempts to copy model capabilities.

The breadth of the cases matters, but the report’s sharper lesson is about method. Anthropic describes actors dividing larger objectives across separate prompts and sessions, making it harder for a model to recognize the harmful purpose of the overall project even when individual requests appear ordinary.

Scale is not limited to a single account or request

Anthropic also says it identified and disrupted illicit model-distillation activity involving seven China-based laboratories since February. Model distillation can be a legitimate training technique, but Anthropic characterizes the activity in these cases as unauthorized extraction of its models’ capabilities.

The company reported that Moonshot AI sent almost 300,000 customer requests to Claude in 10 days. DeepSeek sent more than 12 million exchanges over two weeks in July. Anthropic also attributed more than 151 million exchanges from May through July to an Alibaba-linked campaign targeting Qwen models. It called the campaign its largest measured campaign of this type.

A safety problem built from ordinary-looking pieces

The same pattern appears in the report’s more consequential examples. Anthropic describes AI-assisted development of missile-guidance software in Yemen and autonomous kamikaze-drone software in Russia. In the surveillance case, it says a platform in Mali was designed to monitor roughly 25 million SIM cards across three national mobile networks.

What the report puts in one frame

  • Unauthorized model-copying campaigns can operate at volumes that are difficult to assess one request at a time.
  • Surveillance and weapons-related work can be assembled through AI assistance across multiple tasks.
  • A high refusal rate for explicit malicious prompts does not necessarily describe performance across a fragmented workflow.

That distinction complicates a simple reading of safety results. A refusal rate can describe how a model responds to a direct request. It does not, by itself, show whether the system can connect a chain of related requests into the harmful outcome a user may be pursuing.

Anthropic presents disruption as part of its response, but its example leaves a practical question: Can safeguards spot coordinated misuse when each step looks harmless on its own? That is harder than refusing one dangerous prompt and matters across the different cases the company describes.

Editorial analysis

Our Read

Anthropic’s most important disclosure is not simply that Claude appeared in surveillance and weapons-related work. It is the operational weakness it identifies: a system that rejects an overtly malicious request can still be less dependable when intent is distributed across separate sessions. That shifts the safety challenge from prompt filtering toward recognizing patterns over time, while raising difficult questions about how much context providers should aggregate. The next meaningful evidence will be whether Anthropic can show that its countermeasures catch linked activity without sweeping ordinary use into the same net.

Sources

  1. ndtvprofit.comAnthropic's Misuse Dossier And 2030 Economic Model, Published Days Apart, Show Black Mirror Is Here

Loading discussion...