OpenAI Uses Contractors to Review Real ChatGPT Prompts, 404 Media Finds
The reported Project Lily workflow turns real chats into feedback for future models. Consumer users can opt out, but the setting is default-on and OpenAI says its privacy filter can make mistakes.
Listen to this story
The audio brief
Story brief
3 key pointsA 404 Media investigation details OpenAI’s “Project Lily,” in which hundreds of contractors review real ChatGPT prompts and rate generated replies for model improvement. OpenAI says eligible conversations pass through Privacy Filter, but acknowledges that redaction can miss uncommon or ambiguous identifiers. For consumer accounts, “Improve the model for everyone” is enabled by default; users can opt out, while...
- 01
Contractors reportedly summarize intent, compare multiple answers, score them, and explain failures involving style, flattery, or simulated human experience.
- 02
OpenAI’s Privacy Filter may redact too much or too little when context is limited, according to the company.
- 03
Free, Plus, and Pro accounts have model improvement enabled by default; Enterprise, Business, and Edu disable it by default.
OpenAI is hiring hundreds of contractors to read real ChatGPT prompts and evaluate the chatbot’s replies for model improvement, according to a new 404 Media investigation. The prompts are anonymized, but can still contain sensitive information—putting people behind conversations that users may treat as private.
The outlet reviewed internal instruction guides, Slack channels, real user prompts and the worker rating system. The materials identify the effort only as “Project Lily”; they do not name the OpenAI model being trained or establish whether it is already available to users.
A real prompt becomes a training task
The reported workflow is structured rather than a simple read-through. A contractor reads a user prompt, summarizes the user’s intent, compares multiple generated answers, highlights passages that follow or miss the project’s instructions, scores the responses, and explains the score.
The instructions focus on response quality and style. They direct reviewers to assess whether replies are helpful and natural, while flagging problems such as forced style mimicry, excessive flattery and language that implies the system has human experiences.
Anonymization has limits
Contractors do not see ChatGPT usernames, 404 Media reported. But prompts can still include sensitive or personal details, and a user-memory summary displayed above some prompts can add context about what someone has used the chatbot for and, in some cases, where they may live.
OpenAI says it runs eligible conversations through Privacy Filter before contractor review to detect and remove personal information. The company also acknowledges the system can miss uncommon identifiers or ambiguous private references, and can redact too much or too little when context is limited.
That creates a different privacy question from review of chats flagged for safety or policy concerns. Project Lily, as described by 404 Media, uses real prompts to generate detailed feedback on how well ChatGPT answers and communicates.
The setting determines the boundary
- Free, Plus and Pro accounts have the “Improve the model for everyone” setting enabled by default, 404 Media reported. Turning it off means new conversations stay in chat history but are not used to train ChatGPT.
- Temporary Chats are not used to improve models, OpenAI says. They do not appear in chat history, do not create memories, and are retained for 30 days for safety purposes before deletion.
- Enterprise, Business and Edu customers have model improvement disabled by default, according to 404 Media.
After 404 Media asked whether OpenAI explicitly tells users that humans may read prompts for model improvement, the company added more opt-out detail to its help page. The update still did not explicitly state that humans may read prompts for that purpose, according to the outlet.
Human review is an industry practice, not an OpenAI exception
Anthropic says consumer users can choose to allow their chats and coding sessions to improve its models. When it uses that material, the company says it may include the entire related conversation, custom styles and conversation preferences; it separately says it de-identifies conversations before human review.
The comparison does not resolve Project Lily’s central issue: whether a default-on consumer training setting gives people a clear enough understanding of how their chats may contribute to improvement. OpenAI offers tools to opt out, but the reported contractor workflow makes disclosure—not just data filtering—the point of friction.
Editorial analysis
Our Read
Project Lily makes a routine-sounding data control more consequential. OpenAI has long said users can decide whether conversations improve its models, and it has described safeguards intended to reduce personal information. But the reported use of contractors means that choice can determine whether a real conversation becomes material for human evaluation, not only automated training. The key next test is not whether OpenAI keeps an opt-out: it is whether its consumer-facing disclosures plainly explain the review path and make the choice legible before someone shares a sensitive detail.
Sources
- privacy.claude.comIs my data used for model training? | Anthropic Privacy Center
- openai.comHow ChatGPT learns about the world while protecting privacy
- 404media.coInside ‘Project Lily’: The Humans Reading Your ChatGPT Chats
Loading discussion...
Reader comments
Newest comments first. Replies stay oldest first.