https://www.theregister.com/2025/04/03/openai_copyright_bypass/
Tech textbook tycoon Tim O’Reilly claims OpenAI mined his publishing house’s copyright-protected tomes for training data and fed it all into its top-tier GPT-4o model without permission. This comes as the generative AI upstart faces lawsuits over its use of copyrighted material, allegedly without due consent or compensation, to train its GPT-family of neural networks. OpenAI denies any wrongdoing. O’Reilly (the man) is one of three authors of a study [PDF] titled, “Beyond Public Access in LLM Pre-Training Data: Non-public book content in OpenAI’s Models," issued by the AI Disclosures Project. By non-public, the authors mean books that are available for humans from behind a paywall, and aren’t publicly available to read for free unless you count sites that illegally pirate this kind of material. The trio set out to determine whether GPT-4o had, without the publisher’s permission, ingested 34 copyrighted O’Reilly Media books. To probe the model, which powers the world-famous ChatGPT, they performed so-called DE-COP inference attacks described in this 2024 pre-press paper. …
Regards