
How do people actually use ChatGPT, Claude, and similar tools? The answer available today comes almost entirely from the companies selling those chatbots. That, researchers at Stanford argue, is exactly the problem: without an independent source, there’s no real way to check whether usage reports from OpenAI and Anthropic paint a complete picture or mostly show what looks good. A new project called the “AI Observatory” is trying to close that gap, and its early findings diverge noticeably from the companies’ own accounts.
Key takeaways
- Stanford researchers led by PhD candidate Anka Reuel argue that highly consequential decisions about AI’s benefits and risks are being made on the basis of very limited data curated by the companies themselves.
- Their “AI Observatory” project independently analyzed 24,521 real conversations from 5,000 users across 52 different AI models, spanning 2023 to 2025.
- Applying Anthropic’s own filtering method to that dataset would exclude roughly 48 percent of the conversations from a public report, disproportionately including topics like health, relationships, and sexual content.
- OpenAI’s own September 2025 study shows a similar shift: only about 27 percent of ChatGPT messages are now work-related, down from 47 percent a year earlier.
- Researchers are calling on AI companies to share anonymized raw data with independent scientists instead of publishing only curated summaries.
What the big providers report, and what they leave out
Both Anthropic and OpenAI regularly publish reports on what people use their models for. Anthropic’s Economic Index analyzes millions of Claude conversations but, as Reuel told MIT Technology Review, focuses mainly on productivity-related use in work contexts and largely leaves out personal use. OpenAI went a step further with a September 2025 study conducted with Harvard economist David Deming, analyzing 1.5 million ChatGPT conversations, by the company’s own account the largest study of its kind to date. The result was notable: the share of work-related messages fell from 47 percent a year earlier to roughly 27 percent by June 2025, while personal uses such as writing, seeking information, and everyday practical guidance together made up 78 percent of all interactions. Revealing as those numbers are, they come exclusively from companies’ internal data pools, whose exact collection and filtering methods can’t be fully verified from the outside.
Stanford’s counter-project
That’s precisely where the “AI Observatory” comes in, a project at the Stanford Trustworthy AI Research Lab that Reuel co-leads. Rather than waiting on company data, the team combined seven existing datasets collected with users’ consent: 24,521 conversations spanning 85,633 individual turns, from 5,000 people using 52 different models, including ChatGPT, Claude, Gemini, and Grok, between 2023 and 2025. That’s a considerably smaller dataset than the one to one and a half million conversations Anthropic and OpenAI draw on for their own reports, but it’s one nobody curated in advance. Applying Anthropic’s own filtering logic to this independent dataset would strip out roughly 48 percent of the conversations from a public report. What disappears disproportionately is telling: conversations about health and relationships make up 44.2 percent of the full dataset versus 31.2 percent in the filtered version, sexual content 16.7 percent versus 2.4 percent, and content classified as hateful 27.5 percent versus 5.66 percent. Precisely the areas where users interact with AI systems in the most personal, vulnerable, or sometimes problematic ways are underrepresented in the providers’ official numbers, a pattern already visible in “What ChatGPT Really Knows About You,” which examined the quiet data collection happening in the background. Here the mirror image plays out: what gets left out of the published analyses, deliberately or not.
Why the gap is more than an academic dispute
Reuel puts the consequence bluntly: highly consequential decisions about AI’s benefits and risks are currently being made on the basis of very limited data that nobody outside the companies can really verify. That’s not just academic curiosity. When lawmakers in Brussels or Berlin debate rules for AI chatbots, or when health insurers or schools decide whether to deploy AI assistants, they inevitably lean on exactly the reports whose completeness can barely be checked. David Widder of the University of Texas at Austin proposes replacing separate, company-published reports with something closer to a shared bird’s-eye view, combining data from multiple providers and independent researchers into something comparable. The more obvious, if politically harder, demand is that AI companies share anonymized raw data with independent scientists under strict privacy safeguards, instead of publishing only finished, PR-ready summaries. Until that happens, a basic problem remains: the companies that benefit most from their systems being seen as safe and useful are also the only ones with full access to the data that could actually test that claim.
Where this leaves things
Neither Anthropic’s Economic Index nor OpenAI’s usage study is wrong, both offer real, at times illuminating numbers. The problem lies in what falls between the lines: an entity that can independently verify and supplement those figures. The AI Observatory is still too small to replace the companies’ reports, but large enough to show that the official numbers systematically thin out in specific places, precisely where sensitive, personal use is concerned. For a technology increasingly woven into health, relationships, and everyday decisions, that’s not a footnote. It’s the real gap in the debate.
