Anthropic Opens 750,000 Claude Conversations to Outside Study Without Showing the Chats
The pilot gives outside labs a larger window into real AI use than public chat datasets offer, but Anthropic still runs the analysis, screens the outputs and limits what researchers can inspect.
Listen to this story
The audio brief
Story brief
3 key pointsAnthropic is turning production Claude usage into a research dataset without releasing chat transcripts: Stanford, Oxford and METR analyzed roughly 250,000 conversations each from April–May through company-run classification. The pilot found substantial consequential work and frequent user correction, but its value is constrained by opaque categorization and Anthropic’s review of every cluster, including removals...
- 01
The samples totaled about 750,000 conversations; Claude.ai data excluded Team, Enterprise and API traffic.
- 02
Stanford found 56% of actionable-task conversations involved consequential-or-higher-impact work; 12% were high-stakes.
- 03
Researchers received counts and cluster descriptions, not source chats, limiting their ability to audit classification errors.
Anthropic has created a controlled route for outside researchers to study how people use Claude at production scale. Three groups examined separate samples totaling about 750,000 conversations, but received aggregated results rather than the chats themselves. The arrangement opens a new evidence source for AI-use research while leaving Anthropic at the center of the data pipeline.
Three studies, one protected access model
Stanford’s Social and Language Technologies Lab, Oxford’s Human Information Processing Lab and the evaluation group METR each worked from a separate sample of roughly 250,000 conversations collected in April and May. Stanford and Oxford studied Claude.ai use; METR examined Claude Code sessions. The Claude.ai material covered Free, Pro and Max users, not Team, Enterprise or API traffic.
The distinction is central. Researchers chose analytical questions, known as facets, but Anthropic ran them on its systems. Claude classified conversations, grouped similar responses and returned category counts, percentages and cluster descriptions. The researchers did not view the underlying conversations.
A wider lens, with company-operated guardrails
The groups had publication independence, and Anthropic’s contractual review rights were limited to privacy, policy violations, confidential information and research accuracy. Yet the company manually reviewed every cluster before release. It removed or altered 1.9% of Stanford’s clusters, 3.33% of Oxford’s and 1.8% of METR’s.
Anthropic has published the resulting aggregate outputs under a CC BY 4.0 license on Hugging Face. That makes the released categories available for inspection, but not the source material that produced them.
Stanford’s findings complicate the low-stakes-use narrative
Stanford’s completed study analyzed 249,834 Claude.ai conversations. Among conversations involving an actionable task, 56% involved work classified as consequential or higher-impact: work affecting other people or difficult to reverse. Twelve percent fell into the high-stakes category, including legal and financial guidance.
The collaboration pattern Stanford identified
- Users retained primary responsibility and directed the work in 72% of conversations, with Claude serving in an assisting role.
- Users tended to modify Claude’s output rather than accept it unchanged, especially as the stakes rose.
- Friction appeared in 49.7% of conversations, and users tried to recover in 78.7% of those cases, often by clarifying requests or challenging Claude’s reasoning.
The classification layer is the limiting factor
The privacy design creates an analytical trade-off. Because researchers cannot inspect source chats, poorly phrased facets can generate misleading categories that are hard to detect. Tests on the public WildChat dataset did not always transfer to Claude traffic, and Anthropic removed one Oxford facet after determining its descriptions were unreliable.
Oxford’s study remains underway. Its early patterns link warmer Claude interactions with more positive user behavior, while refusals and disagreement coincide with user pushback; those are associations, not evidence that one response style caused the other. METR’s unfinished coding study has preliminary indications that newer Claude models deliver significant speedups over older versions, using model estimates of no-AI task time as part of its method.
Privacy safeguards do not erase disclosure risk
A privacy red team at Imperial College London reported that it could not reidentify users from released material. It did, however, connect one cluster to a widely used open-source project through distinctive language. Anthropic says future studies will use higher minimum-user thresholds and less distinctive cluster descriptions.
Anthropic is soliciting proposals for future access, while warning that privacy and review requirements will make expansion slow. The pilot’s value will rest on whether that controlled path can support more varied questions without weakening the protections that keep raw conversations private.
Editorial analysis
Our Read
Our read: this is a more meaningful opening than publishing a lab-authored usage report, because Stanford, Oxford and METR set their own questions and retain control of their conclusions. But it is not independent access to production data in the conventional sense: Anthropic selects samples, operates the classification system and reviews released clusters. The next test is whether future projects can produce findings that meaningfully challenge Anthropic’s own priorities or framing while still surviving those privacy controls. The program’s slow scaling warning makes that question operational, not merely philosophical.