Anthropic has released findings from a pilot program that allowed three outside research groups to study aggregated data from roughly 250,000 Claude and Claude Code conversations. One study found that users delegate consequential tasks to the AI more often than earlier research had suggested, especially when seeking professional legal or financial guidance.
Kunal Handa and colleagues announce for Anthropic that the company gave researchers access through Anthropic Insights, a privacy-preserving analysis system formerly called Clio. The researchers did not receive raw chat transcripts. Instead, they designed questions that Claude applied across large sets of conversations, then reviewed only aggregated results.
The pilot involved Stanford University’s Social and Language Technologies Lab, the University of Oxford’s Human Information Processing Lab, and the nonprofit model evaluation group METR. Each team developed its own research questions and conducted its own analysis. Anthropic says its ability to review findings was limited to privacy, security, confidential information, accuracy and material that could help users evade safety rules.
High stakes work reaches AI tools
Stanford’s SALT Lab examined how people work with Claude and where that collaboration fails. Its early findings challenge the assumption that users reserve AI for low risk tasks.
More than half of the conversations studied involved users delegating work described as consequential. These are tasks that may affect other people or be difficult to reverse. Users were particularly likely to bring such work to Claude when they wanted professional guidance on legal or financial issues.
The researchers also found that users generally remain involved. In nearly three quarters of conversations, people set the direction and Claude played a supporting role. Users often adapted the model’s output rather than copying it directly.
Friction was common, according to the study, but it was not always harmful. Users frequently had to clarify requests, identify misunderstandings and refine instructions. That process could improve the final result and keep users engaged with the underlying problem.
Emotion, behavior and productivity
Oxford’s Human Information Processing Lab is investigating how users feel during AI interactions and how those feelings relate to Claude’s responses. Its preliminary analysis links warmer responses from Claude with more positive user behavior. Refusals or disagreement often appeared alongside users pushing back, while more unusual responses coincided with greater intellectual engagement.
The Oxford team also reports that patterns of frustration, enjoyment and absorption in Claude conversations resemble patterns observed in a separate study of general web browsing. The researchers have not yet published their full write-up.
METR is studying whether coding agents create measurable productivity gains. Its work remains in progress, but early results suggest that newer Claude models may save users more time than earlier versions. METR compared Claude’s estimates of how long tasks would have taken without AI with task duration estimates from a prior developer study. According to Anthropic, those estimates showed a meaningful correlation with actual completion times.
Anthropic says the program revealed practical limits to scaling independent access. The system depends on carefully worded questions, because vague prompts can create misleading categories. External teams tested questions with the public WildChat dataset, but its more casual content did not always reflect Claude’s real world usage.
The company also removed or altered a small number of categories that could reveal ways to bypass safeguards. Anthropic says these changes affected fewer than 5 percent of categories and conversations in each study. It is now collecting expressions of interest from researchers while assessing whether it can expand the program without weakening privacy, safety or research quality.
Stay up to date
AI for content creation: the latest tools, tips and trends. Every two weeks in your inbox: