A proposed class action filed in the U.S. District Court for the Northern District of California accuses OpenAI of routing real ChatGPT conversations to outside contractors without adequate disclosure. The lawsuit centers on "Project Lily," an internal program where contractors hired through third-party staffing firms summarize user prompts and score model responses on a one-to-seven scale. Although OpenAI employs automated filtering before human review, the complaint alleges that personal details sometimes reach these reviewers via dashboards that include "user memories summaries" revealing location, profession, or personal life details. The plaintiffs argue that while OpenAI’s privacy policy mentions some external data access, it fails to specifically flag data annotation or human evaluation vendors.
The legal filing includes eight claims under California’s Unfair Competition Law, Consumer Privacy Act, and common-law intrusion upon seclusion. With ChatGPT serving more than 900 million weekly users, many of whom input sensitive health, legal, or financial information, the suit seeks damages and specific injunctive relief. Plaintiffs demand that OpenAI require opt-in consent before human review, switch the "Improve the model for everyone" setting to off by default, add clear warnings within the chat interface, and potentially delete work product tied to reviewed conversations. OpenAI was served on September 2 and has until October 13 to respond.
This litigation highlights the friction between large-scale AI model training methodologies and consumer privacy expectations. While reinforcement learning from human feedback is standard industry practice for improving model accuracy and reducing sycophancy, the lack of explicit, granular consent for human review creates significant regulatory exposure. The case underscores how existing privacy frameworks, such as California’s Consumer Privacy Act, may not adequately address the nuances of modern AI data pipelines where automated filters fail to catch all personally identifiable information before it reaches human annotators.
For institutional adoption and market structure, this development signals growing scrutiny on data governance practices within major AI providers. If courts enforce strict opt-in requirements or mandate default-off settings for data sharing, it could increase operational costs and slow down model iteration cycles. The outcome will likely influence how other tech companies structure their contractor agreements and user disclosures, potentially leading to a bifurcation in service tiers where premium, privacy-focused offerings command higher prices compared to free models reliant on broad data usage rights.


