AI Copyright Case: Execs Knew Training Data Was 'Pirated,' Filings Allege
Unsealed Filings Expose AI Giants' 'Lawless' Training Tactics
Newly unsealed court documents in the landmark copyright case Authors Guild v. OpenAI have thrust the internal knowledge and alleged intentional misconduct of Microsoft and OpenAI executives into the spotlight. The filings, made public on September 17, 2026, paint a damning picture of decision-making at the highest levels, suggesting that key players were fully aware they were using pirated books to train their AI models—and proceeded anyway.
The documents, filed as part of a motion for partial summary judgment, quote internal communications where executives allegedly described the practice as using books from a “sketchy Russian website” and acknowledged the technology would “make people unemployed.” The revelations add a new layer of legal exposure for the tech giants, shifting the narrative from accidental infringement to deliberate, knowing misconduct.
Executives' Own Words: 'Pirated Stuff' and 'Existential Threat'
Among the most explosive allegations is a quote attributed to a top AI executive: “We trained GPT-3 with pirated stuff!” This statement, along with a 2023 memo from Microsoft Director of Applied Science Brent Hecht, forms the crux of the plaintiffs' argument. Hecht allegedly described the mass scraping of online content as an “astonishing theft” and potentially the “largest theft of labour in human history,” a characterization that has since reverberated across the tech industry.
The briefs also highlight a warning from within OpenAI itself. Nick Turley, OpenAI’s head of ChatGPT, reportedly described the company’s products as “largely substitutive” for the original works, with that substitutability expected to grow as the technology improves. The plaintiffs argue this is direct evidence that the companies understood their products would cannibalize the market for human-authored books, posing what they call an “existential threat to those who write and publish books.”
The 'Doom Loop' and a Culture of Recklessness
Perhaps most striking is the internal acknowledgment of a “doom loop” of AI content strategy, a term reportedly used by Microsoft itself in the unsealed documents. This concept describes a vicious cycle where AI systems consume human-generated content, degrade the quality of the information ecosystem, and then require even more fresh human content to retrain—leading to further degradation. Despite this awareness, the filings allege, the companies pressed forward with what they knew was illegal activity.
Authors Guild CEO Mary Rasenberger did not mince words in response to the filings. “These filings reveal shocking disdain for writers and their work through repeated, intentional decisions to steal books rather than pay for them,” she said. “OpenAI pursued its mass piracy scheme even in the face of clear evidence that doing so would degrade American culture by substituting human works with AI slop and would put thousands of writers out of work.”
Legal Context: Fair Use vs. Deliberate Infringement
OpenAI and Microsoft have consistently defended their practices under the “fair use” doctrine, arguing that training on publicly available text is transformative and does not serve as an unlawful market substitute. However, the new evidence of alleged knowing misconduct could severely undermine that defense. Courts often consider the “good faith” of the user in fair use analyses, and internal communications suggesting an awareness of illegality could be pivotal.
The case, which is part of multidistrict litigation in the Southern District of New York, includes plaintiffs such as John Grisham, George R.R. Martin, Jodi Picoult, and Jonathan Franzen. It is being led by Susman Godfrey’s Justin A. Nelson. A hearing is expected in early 2027, with more briefing anticipated in the coming months.
Broader Implications for the AI Industry
Beyond the immediate legal battle, these filings have profound implications for the entire AI ecosystem. They expose a systemic reliance on copyrighted material that many creators have long suspected but could not prove. The documents also raise critical questions about the sustainability of the internet’s economic model, as AI systems increasingly provide answers without driving traffic to original sources.
For the tech industry, the case represents a potential inflection point. If the court rules against Microsoft and OpenAI, it could force a fundamental shift in how training data is sourced, potentially requiring licensing agreements and fair compensation for creators. This would have a ripple effect on every AI company that has built products on the backs of human labor.
The unsealed briefs also come on the heels of similar revelations in a parallel case brought by news media organizations, suggesting a coordinated legal assault on Big AI’s content practices. As the litigation progresses, the court’s decisions could set precedents that define the boundaries of AI training for decades to come.
What's Next?
The immediate next steps involve further briefing and a potential hearing in early 2027. The court will need to weigh the voluminous evidence of alleged misconduct against the companies’ formal legal arguments. Regardless of the outcome, the unsealed documents have already achieved a significant victory for the plaintiffs: they have shifted public perception and put the tech giants on the defensive.
For authors and publishers, the case is more than a legal battle—it is a fight for the future of their profession. As one internal Microsoft memo reportedly put it, the company was aware it was engaging in a “doom loop” that would destroy the very ecosystem it was feeding on. The question now is whether the courts will hold them accountable for that knowledge.
Related News

Parley: Federated IRC Chat That Speaks Plain Protocol

Ollaya: Run Jev-Style Decision Models Locally at Millisecond Speeds

How OpenAI Agents Hacked Hugging Face: New Details Revealed

Claude Opus 5.5 Turns Code Into Studio-Quality Explainer Videos

DHH Declares 'Pencils Down' on Hand-Written Code in Rails World 2026 Keynote

