On September 14, 404 Media’s Joseph Cox published an account of an internal OpenAI program called Project Lily. Hundreds of contractors are employed to read real ChatGPT conversations and rate the model’s responses — part of the ordinary machinery of reinforcement learning from human feedback, aimed at specific behavioral problems including the model anthropomorphizing itself and agreeing with users too readily.
OpenAI says identifying information is stripped before prompts reach reviewers, and that reviewers do not see usernames. It also acknowledges that sensitive details can still get through. Anthropic confirmed it uses human review as well. So, in various forms, does every major lab.
One contractor’s observation is the line that traveled: users, he said, do not imagine “some contractor somewhere is analyzing the conversations.”
He is almost certainly right, and that gap — between what is disclosed and what is understood — is the actual story. The practice is neither secret nor novel. It is disclosed in privacy policies, it is standard across the industry, and it is how these systems are improved. The surprise is not evidence of deception. It is evidence that the disclosure regime we have built does not produce comprehension, and a decade of consent-banner jurisprudence has trained everyone involved to treat those as the same thing.
What Makes This Different From Alexa
Human review of voice assistant recordings produced a wave of coverage in 2019 and a set of regulatory consequences that shaped the sector. The obvious response to Project Lily is that this is the same story with a new logo.
It is not, and the difference is in what people say to the two products.
A voice assistant transcript is overwhelmingly banal: timers, weather, music, the occasional accidental recording. The privacy harm was real, but it was mostly a harm of ambient surveillance — the sense of being overheard rather than the exposure of anything specific.
A ChatGPT conversation is a different artifact. People use these systems as therapist, lawyer, doctor, confessor, and career counselor. They paste in medical results, contract drafts, severance agreements, financial statements, and the kind of personal material that people historically disclosed only under professional privilege. They do it because the interface is private, the response is non-judgmental, and there is no visible second party.
“The product is doing exactly what it was designed to do, and that is the problem,” says Hassan Taher, an AI analyst and author who advises organizations on enterprise AI strategy. “Every design decision in a chat assistant is aimed at lowering the barrier to disclosure. It’s conversational, it’s patient, it doesn’t visibly judge, and it responds better the more context you give it. That’s not a manipulation — it’s genuinely how you get a useful answer. But the effect is that you have built an instrument that elicits confession, and then applied a data governance model designed for search logs. Those two things were never going to fit.”
The Anonymization Problem Is Not a Tooling Problem
OpenAI’s position is that a filter removes identifying information before review. The honest assessment is that this filter cannot work reliably, and the reason is structural rather than a matter of engineering investment.
Redaction systems find named entities: people, places, phone numbers, account identifiers, dates. They are quite good at that now. What they cannot do is recognize when a combination of unremarkable details resolves to one person. A message describing a specific medical condition, a job title, a company size, and a city contains no personally identifying information by any conventional definition, and identifies exactly one human being.
Long conversations make this worse rather than better. A single message may be genuinely anonymous. A forty-turn thread in which the user has described their team, their dispute with a named counterparty, their diagnosis, and their timeline is a dossier, and no entity-recognition pass changes that. The same asymmetry runs through the ethics of how these systems are trained, which Hassan Taher has addressed in examining the relationship between AI training and the right to privacy — the material that makes a model better is precisely the material that is most costly to expose.
“Anonymization was invented for structured records, where you could remove a column and mean it,” Taher notes. “Free-text conversation has no columns. There is no field called ‘identity’ that you delete. Identity is distributed across the whole document in a way that is obvious to any human reader and invisible to any classifier. Companies deploying these systems should stop treating redaction as a control and start treating it as a courtesy. It reduces casual recognition. It does not prevent determined identification, and it should never be the thing standing between a customer’s data and a third-party reviewer.”
Where the Enterprise Exposure Actually Sits
For businesses, the consumer privacy angle is the less consequential half of this.
The more consequential half is that a substantial volume of corporate material flows through consumer AI accounts every day. Employees paste in contracts, strategy documents, customer lists, unreleased financials, and code — on personal logins, outside any enterprise agreement, under consumer terms that permit exactly the kind of human review Project Lily describes.
The enterprise tiers of these products carry different terms. Zero-retention arrangements exist. Contractual commitments against human review and training use are available and increasingly standard. None of that protects an organization whose employees are using the free tier on a personal account, because the enterprise agreement governs the enterprise account and nothing else.
This is a governance failure with an unusually clear remedy, and it is the same failure that shows up in most of the research on why AI deployments underperform — a gap between what leadership authorized and what the organization actually does, which Hassan Taher has examined in the MIT finding that the overwhelming majority of AI projects fail to deliver measurable return.
Four things are worth doing, and none of them are difficult:
- Determine whether your enterprise agreement actually excludes human review. Many organizations assume it does because they are paying. Data processing terms vary considerably by tier and by vendor, and the relevant clause is often in a linked appendix rather than the master agreement.
- Find out what fraction of AI usage in your organization is happening on personal accounts. For most companies that have not deliberately measured this, the answer is a majority. Network telemetry will tell you in an afternoon.
- Give people a sanctioned tool that is good enough to use. Shadow AI usage is almost always a response to a sanctioned option that is slower, worse, or gated behind an approval process. Policy does not beat convenience. A better default does.
- Say out loud that anything typed into a consumer AI product may be read by a person. Not as a legal disclaimer — as a plain operational fact, delivered the way you would tell someone not to email a spreadsheet of customer records to their personal address.
The Broader Point
There is a version of this story in which OpenAI is the villain. It is not a well-supported version. The company disclosed the practice, built a filter, and is doing the thing that is required to make these systems less harmful — including, specifically, training the model to stop pretending to be a person and to stop telling users what they want to hear. Those are improvements that only happen because humans read the failures.
The uncomfortable conclusion is that the safety work and the privacy cost are the same activity. A model cannot be tuned away from sycophancy by a system that has never seen sycophancy in the wild. The material that makes these products better is, unavoidably, the material users would least like reviewed.
“We are going to have to get more honest about that trade rather than pretending it can be engineered away,” Taher says. “The right response is not outrage that humans are involved, because humans being involved is what makes the system safer. The right response is to insist on the specifics — how long the data is held, who can see it, under what contractual controls, with what audit trail, and whether a user can opt out without losing the product. Those are answerable questions. Most vendors currently answer them in a privacy policy nobody finishes reading, and most buyers currently accept that. Both halves of that arrangement need to change.”
