
How people actually use ChatGPT, Claude, or Gemini is known above all to the companies behind those services. They now publish remarkably detailed usage reports. Yet every one of those statistics inevitably shows only one slice: one platform, its users, and its measurement method. The new AI Observatory aims to close that gap. The public research platform brings anonymized conversation data from seven sources together under a shared taxonomy. It is not a census of AI use. But it is an important counterweight to the habit of explaining AI’s social impact solely through vendor numbers.
Key takeaways
- The AI Observatory combines samples of real chatbot conversations from seven datasets and more than 52 models.
- The platform classifies conversations using 145 features and exposes differences among sources instead of smoothing them into one seemingly perfect number.
- The data assessed so far argues against treating AI only as a tool for office work and coding: a substantial share of use is personal, advisory, or emotional.
- Even an independent dataset is selective. Public chat data and voluntarily contributed conversations are not the same as a representative survey of all users.
A window into an otherwise closed world
The AI Observatory is an interactive companion to academic work on AI adoption and capabilities. According to the project, it aggregates seven sources with conversation samples dated from 2023 through 2026. These include publicly available conversation datasets such as WildChat as well as different model and platform sources. Each conversation is described with a shared taxonomy of 145 properties, from prompt structure and response role to the course of the conversation and its task. The platform’s strength is not that it settles every question. Its value is that data sources can be viewed alongside each other.
That comparison is often what is missing. OpenAI, Anthropic, and Google can access very large volumes of their own usage data, but they must account for privacy, business interests, and their own product logic. OpenAI published data this summer on ChatGPT’s growing and broadening use. Anthropic recently expanded its Economic Index with more granular time series, separate analyses for conversations and API use, and surveys. Those are valuable contributions. They do not automatically explain how users behave outside the respective platform or whether a classification would look the same at a competitor.
The difference resembles the gap between a mobile provider’s customer-satisfaction study and a comparative consumer study. The first can be careful and still only observe its own customers. The Observatory does not replace vendor reports. It creates a shared layer on which their results can be checked, contextualized, and at times challenged.
Private use disappears in the workplace debate
Looking across several sources is especially revealing because it changes the usual labor-market discussion. The Washington Post reported on an analysis of more than 23,000 chatbot conversations from 2023 to 2025. Forty-eight percent of the conversations were not work related. The researchers also observed a growing number of conversations about health and relationships. That does not make chatbots reliable therapists or medical advisers. It does mean that debates about productivity, job cuts, and software development capture only part of the real points of contact.
That changes the question of risk as well. When AI is used to write code, errors are often technically testable and responsibilities are comparatively clear. When people seek advice about health, conflict, or loneliness, different questions move to the foreground: How clearly does a system state its limits? What data is created? And when should an answer do more than sound plausible and point to professional help? Such use cannot be sensibly inferred from enterprise or API figures alone.
The recent debate over AI measurement shows the same point in another area. Our report on Google’s double-blind tests argued that strong benchmark scores are no substitute for robust real-world evaluation. The same thought applies to usage research: a large number of prompts is not yet an explanation of the place a system holds in everyday life.
Independent does not mean representative
The most important caveat should stay in view when reading the charts. A dataset made from publicly shared, exported, or voluntarily provided chats does not automatically represent everyone who uses AI. Particular languages, countries, occupations, and platforms may be overrepresented or underrepresented. Models differ sharply as well: short, template-like requests to older systems tell a different story from long, iterative conversations with newer models. The Observatory makes many of those differences visible. It does not make them disappear.
Classification itself is another limitation. Categories such as work, education, advice, or personal use help comparison, but they remain interpretations. A conversation about a job application may be career advice, writing assistance, and a personal concern at the same time. That is why it is useful that the platform does not merely produce a headline, but discloses data origin, time period, and features. Anyone drawing policy or economic conclusions should always ask: Which conversations were captured, which are missing, and how were they counted?
That transparency is more productive than the dream of one neutral AI metric. Different measurements can sit next to one another when their limits are clear. Vendors can provide large and current data. Public research can compare methods and identify blind spots. Surveys can capture motives that conversation records do not show. Taken together, they create a more durable picture than any single dashboard.
What policymakers, researchers, and users should do with this
For policymakers and regulators, the platform signals that AI impact assessment needs data access without exposing private conversations. Aggregated, privacy-preserving interfaces and accountable research programs would be more useful than blanket demands for raw data. Anthropic is already testing one route through its Insights program, in which external groups can pose their own research questions to aggregated usage data. The key is for such offerings not to stop at a single provider and to permit independent methods.
For readers, the lesson is more modest but useful: neither an AI vendor’s success announcement nor a single chart can fully describe reality. When you see a statistic about AI and work, ask about the platform, period, sample, and definitions. This is not academic hair-splitting. The answers determine whether schools, employers, and lawmakers are discussing actual uses or merely what is easiest to count.
The AI Observatory is therefore above all a starting point. It does not prove how people use AI. It shows that the question is larger than any individual company’s data store. As AI enters more private and professional routines, a public measurement practice that brings sources together, makes uncertainty visible, and discloses its own limits will become more important. That is less spectacular than a new model. For sensible AI policy, it may matter more in the long run.
