
Real AI usage data usually lives inside a few labs. On August 26 Anthropic said a spring pilot let three outside groups pose their own questions through Anthropic Insights (formerly Clio) on roughly 250,000 Claude.ai / Claude Code conversations from April–May 2026. Officially, this is the first time external researchers published independent studies on an AI company’s own usage data; contract review covered privacy, misuse-bypass detail, confidential info, and accuracy — partners could publish even when findings were inconvenient for Anthropic.
[1]Three early threads share one post. Stanford SALT: over half of conversations involved consequential tasks (affect others or hard to undo), especially legal/financial guidance; in nearly three-quarters, humans set direction and usually edited rather than pasted outputs, with friction often clarifying intent. Oxford: warmth, refusal, eccentricity, and plain help from Claude co-occurred with user positivity, pushback, intellectual engagement, and satisfaction — emotional patterns resembled everyday web browsing. METR: comparing Claude’s “time without AI” estimates to actual durations tentatively shows newer models save more time, and estimates correlate with a prior developer study. Researchers never saw raw chats — only aggregated categories and shares.
[1]The sharp edge is institutional, not another dashboard. External work usually eats a lab’s own narrative or casual public corpora; Insights tries a third path: outside questions, inside privacy pipelines, outside write-ups. Costs are real: wording-sensitive classifiers, no week-long iteration on raw text, WildChat-tuned categories that misread Claude traffic, and misuse clusters describing how safeguards were bypassed rewritten or removed (<5% of categories/conversations per study). Author’s take: oversight needs publishable question rights, not more in-house white papers.
[1]For policy, the pilot shows privacy-preserving aggregation plus external questions can work — and that scaling is slow and expensive. For other labs, the “publish even if inconvenient” clause becomes a benchmark. For product teams, SALT’s “over half consequential” undercuts the soothing story that users only draft low-stakes copy. An expression-of-interest form signals appetite to expand without promising throughput.
[1]If alignment evals are exams labs give models, the Insights pilot is a spot-check society gives labs. The sample is still small and some Oxford/METR write-ups are unfinished, but the direction is right: whoever owns the question list owns the “social impact of AI” story. Next contest: scale without losing the right to publish inconvenience.
[1]