
Anyone typing a question into ChatGPT usually assumes only a machine reads it. A new investigation by 404 Media shows otherwise: behind the scenes, hundreds of external contractors read real conversations to rate the model’s answers. The internal name for the program is “Project Lily,” and it potentially affects millions of users who have no idea it exists.
Key takeaways
- On September 14, 2026, 404 Media revealed that OpenAI runs an internal program called “Project Lily,” in which hundreds of external contractors read and rate real ChatGPT prompts.
- Reviewers see four AI-generated answer variants per request and score them on a scale of 1 to 7.
- Recruitment runs through a firm called Crossing Hurdles, with pay processed via the staffing platform Mercor; one contractor told 404 Media they earned more than $50 an hour.
- For ChatGPT Free, Plus, and Pro accounts, data use for model improvement is on by default (opt-out); Anthropic says it requires an active opt-in.
- Users can turn off the setting labeled “Improve the model for everyone,” though OpenAI still retains new chats for up to 30 days for abuse monitoring even after opting out.
How Project Lily works
According to internal documents reviewed by 404 Media reporter Joseph Cox, reviewers are shown four different ChatGPT-generated answer drafts for each real user prompt. They rate the drafts on a scale of 1 to 7, flag problematic passages, and summarize the original request in their own words. One stated goal of the program is to make ChatGPT less sycophantic and less anthropomorphic — dialing back the tendency toward exaggerated agreement that OpenAI itself has repeatedly acknowledged as a problem. An upstream “privacy filter” model is supposed to strip names and other personal details before a prompt reaches a human reviewer. But OpenAI acknowledged to 404 Media that sensitive details can still slip through — for instance, via automatically generated memory summaries that can hint at a user’s prior conversations or approximate location.
What OpenAI’s own rules say
The key toggle in ChatGPT’s settings is called “Improve the model for everyone,” found under Data Controls. It is on by default for personal Free, Plus, and Pro accounts, while Enterprise, Business, Team, and Edu accounts have training data use off by default. Turning the toggle off, according to OpenAI, stops new conversations from feeding into training — previously used chats are unaffected. Even after opting out, though, OpenAI still stores new conversations for up to 30 days for abuse review before deleting them. What remains unclear from current reporting is whether this toggle actually prevents an individual chat from being seen by a Project Lily reviewer, or whether it only affects automated training. OpenAI has not publicly clarified that distinction.
A case for GDPR?
No European or German data protection authority has issued an immediate response to the revelation so far — the investigation is still fresh, and that silence shouldn’t be read as a clean bill of health. Some journalistic analyses already point to the EU Court of Justice’s reasoning in EDPS v SRB, which held that data controllers generally must inform individuals before sharing their data with third parties, regardless of whether the receiving party can actually identify the person. Whether that reasoning applies to Project Lily is an interpretation by outside observers, not a regulatory ruling. Historically, OpenAI has already felt Europe’s scrutiny once: Italy’s data protection authority, Garante, fined the company 15 million euros in December 2024 over a lack of legal basis for data processing and insufficient age verification. That case has no direct connection to Project Lily, but it shows European regulators are watching ChatGPT closely.
How other providers handle human review
Anthropic also confirmed to 404 Media that it uses human feedback to improve its models — but says it does so only for users who actively opt in, with account identifiers such as email addresses stripped beforehand. At Google, according to Google’s own support pages, specially trained teams review conversations when a user’s “Gemini Apps Activity” is turned on; that can be switched off in Google account settings. The key difference from OpenAI’s approach isn’t the principle of human quality review itself — every major provider now uses it — but rather opt-in versus opt-out, and how clearly users are informed beforehand. That’s exactly where the criticism of Project Lily lands: according to 404 Media, most of ChatGPT’s more than 900 million users worldwide likely have no idea their words are shown to real people.
Conclusion
That language models need human feedback to get better and safer is not new — RLHF has been an industry standard for years, as the debate around OpenAI’s and Anthropic’s recent safety-pacing plans also shows. What’s striking about Project Lily is something else: the scale, the concrete payment chain running through Crossing Hurdles and Mercor, and above all the lack of transparency toward the people actually affected. Anyone who doesn’t want a stranger reading their prompts should check their data controls now and think twice before typing especially sensitive requests — much like tools that already reveal what ChatGPT really knows about its users suggest. Whether regulators will find the practice legally objectionable remains open based on current reporting. That they’ll be taking a closer look seems likely after this report.
Sources
- 404 Media: Inside 'Project Lily': The Humans Reading Your ChatGPT Chats
- Golem: Datenschutz bei KI: OpenAI lässt KI-Prompts von Menschen lesen
- The Next Web: Hundreds of contractors are reportedly reading real ChatGPT conversations
- Tom's Guide: OpenAI paid contractors to read ChatGPT conversations – here's how to protect yourself

