Toolspublished

Anthropic Opens 750,000 Claude Conversations to Outside Study Without Showing the Chats

The pilot gives outside labs a larger window into real AI use than public chat datasets offer, but Anthropic still runs the analysis, screens the outputs and limits what researchers can inspect.

By 4 min read
Anthropic Opens 750,000 Claude Conversations to Outside Study Without Showing the Chats

Listen to this story

The audio brief

About 1:48
0:001:48
Read transcript
Anthropic has given outside researchers a look at three-quarters of a million real Claude conversations—without showing them the conversations themselves. The pilot involved Stanford, Oxford, and METR, each studying a separate sample of roughly 250,000 sessions collected in April and May. Stanford and Oxford examined Claude.ai use among Free, Pro, and Max users. METR studied Claude Code sessions. Team, Enterprise, and API traffic was excluded. The researchers chose the questions, or “facets,” but Anthropic ran the analysis. Claude classified the conversations, grouped similar responses, and returned counts and descriptions. That protects privacy, but it also means researchers cannot audit the source material when a category looks wrong. Anthropic manually reviewed every cluster, removing or changing between 1.8 and 3.33 percent of the groups before publication. Stanford’s results suggest Claude is often used for consequential work, not just casual prompting. Of conversations involving an actionable task, 56 percent concerned work affecting other people or difficult-to-reverse outcomes, and 12 percent were high-stakes, including legal and financial guidance. Users directed the work in 72 percent of cases, and friction appeared in nearly half—yet users often tried to recover by clarifying or challenging Claude. A privacy red team at Imperial College London found no user reidentification, but it did link one cluster to an open-source project through distinctive wording. Anthropic says future studies will use larger minimum-user thresholds and less distinctive descriptions. The key constraint is whether this company-run pipeline can broaden research without exposing the chats or making the classifications impossible to verify.

Story brief

3 key points

Anthropic is turning production Claude usage into a research dataset without releasing chat transcripts: Stanford, Oxford and METR analyzed roughly 250,000 conversations each from April–May through company-run classification. The pilot found substantial consequential work and frequent user correction, but its value is constrained by opaque categorization and Anthropic’s review of every cluster, including removals...

  1. 01

    The samples totaled about 750,000 conversations; Claude.ai data excluded Team, Enterprise and API traffic.

  2. 02

    Stanford found 56% of actionable-task conversations involved consequential-or-higher-impact work; 12% were high-stakes.

  3. 03

    Researchers received counts and cluster descriptions, not source chats, limiting their ability to audit classification errors.

Anthropic has created a controlled route for outside researchers to study how people use Claude at production scale. Three groups examined separate samples totaling about 750,000 conversations, but received aggregated results rather than the chats themselves. The arrangement opens a new evidence source for AI-use research while leaving Anthropic at the center of the data pipeline.

Three studies, one protected access model

Stanford’s Social and Language Technologies Lab, Oxford’s Human Information Processing Lab and the evaluation group METR each worked from a separate sample of roughly 250,000 conversations collected in April and May. Stanford and Oxford studied Claude.ai use; METR examined Claude Code sessions. The Claude.ai material covered Free, Pro and Max users, not Team, Enterprise or API traffic.

The distinction is central. Researchers chose analytical questions, known as facets, but Anthropic ran them on its systems. Claude classified conversations, grouped similar responses and returned category counts, percentages and cluster descriptions. The researchers did not view the underlying conversations.

A wider lens, with company-operated guardrails

The groups had publication independence, and Anthropic’s contractual review rights were limited to privacy, policy violations, confidential information and research accuracy. Yet the company manually reviewed every cluster before release. It removed or altered 1.9% of Stanford’s clusters, 3.33% of Oxford’s and 1.8% of METR’s.

Anthropic has published the resulting aggregate outputs under a CC BY 4.0 license on Hugging Face. That makes the released categories available for inspection, but not the source material that produced them.

Stanford’s findings complicate the low-stakes-use narrative

Stanford’s completed study analyzed 249,834 Claude.ai conversations. Among conversations involving an actionable task, 56% involved work classified as consequential or higher-impact: work affecting other people or difficult to reverse. Twelve percent fell into the high-stakes category, including legal and financial guidance.

The collaboration pattern Stanford identified

  • Users retained primary responsibility and directed the work in 72% of conversations, with Claude serving in an assisting role.
  • Users tended to modify Claude’s output rather than accept it unchanged, especially as the stakes rose.
  • Friction appeared in 49.7% of conversations, and users tried to recover in 78.7% of those cases, often by clarifying requests or challenging Claude’s reasoning.

The classification layer is the limiting factor

The privacy design creates an analytical trade-off. Because researchers cannot inspect source chats, poorly phrased facets can generate misleading categories that are hard to detect. Tests on the public WildChat dataset did not always transfer to Claude traffic, and Anthropic removed one Oxford facet after determining its descriptions were unreliable.

Oxford’s study remains underway. Its early patterns link warmer Claude interactions with more positive user behavior, while refusals and disagreement coincide with user pushback; those are associations, not evidence that one response style caused the other. METR’s unfinished coding study has preliminary indications that newer Claude models deliver significant speedups over older versions, using model estimates of no-AI task time as part of its method.

Privacy safeguards do not erase disclosure risk

A privacy red team at Imperial College London reported that it could not reidentify users from released material. It did, however, connect one cluster to a widely used open-source project through distinctive language. Anthropic says future studies will use higher minimum-user thresholds and less distinctive cluster descriptions.

Anthropic is soliciting proposals for future access, while warning that privacy and review requirements will make expansion slow. The pilot’s value will rest on whether that controlled path can support more varied questions without weakening the protections that keep raw conversations private.

Editorial analysis

Our Read

Our read: this is a more meaningful opening than publishing a lab-authored usage report, because Stanford, Oxford and METR set their own questions and retain control of their conclusions. But it is not independent access to production data in the conventional sense: Anthropic selects samples, operates the classification system and reviews released clusters. The next test is whether future projects can produce findings that meaningfully challenge Anthropic’s own priorities or framing while still surviving those privacy controls. The program’s slow scaling warning makes that question operational, not merely philosophical.