TechForge

October 22, 2024

  • PRH updates its copyright notice to ban AI training on its books.
  • Aims to protect authors’ intellectual property.

Penguin Random House (PRH) has taken a firm position opposing the use of its books’ content to train AI systems, a watershed moment in the publishing industry’s response to tech developments. According to The Bookseller, PRH has updated the copyright page of both new and reprinted books to specifically stop AIs using its works for training.

The updated copyright statement includes the phrase: “No part of this book may be used or reproduced in any manner for the purpose of training artificial intelligence technologies or systems.” Additionally, PRH notes that it “expressly reserves this work from the text and data mining exception,” in alignment with European Union laws on the matter.

The step makes PRH one of the first major publishers to publicly address the increasing use of copyrighted works to train AI models, such as those employed by tech companies for chatbots and other digital tools. As AI continues to evolve and integrate into various sectors, publishers are grappling with how to protect their content from being exploited by such systems.

Preemptive action in a shifting legal landscape

PRH’s move follows a series of copyright infringement cases in the US, where reports have surfaced of large volumes of pirated books being used by AI firms to train their models. With the legal landscape around AI still evolving, PRH’s decision signals an effort to preemptively safeguard its intellectual property.

Although the language printed on the copyright page of PRH books may appear as a legal defence, it does not constitute a direct change to copyright law. In fact, the new clause functions similarly to a “robots.txt” file, which many websites use to signal that they do not want their content scraped by AI systems or web crawlers. While this serves as a formal request, it lacks legal enforceability – compliance is voluntary.

The same is true for PRH’s new copyright notice; copyright protections exist regardless of the notice, and legal defences such as ‘fair use’ remain, whether or not they are mentioned in the book’s copyright page.

A growing dilemma between AI and copyright

In recent years, the emergence of large language models (LLMs) has brought data usage to the limelight. Models including well-known tools such as OpenAI’s ChatGPT and Google’s Bard are trained on massive datasets sourced primarily from publicly-available text and, in some cases, copyrighted materials. This creates a quandary for publishers: on the one hand, they want to protect their intellectual property while also adapting to an increasingly digital and AI-driven world.

The importance of this debate is underscored by PRH’s public statement in August 2023, in which the company vowed to “vigorously defend the intellectual property that belongs to our authors and artists.” The sentiment is echoed by the Authors’ Licensing and Collecting Society (ALCS), which has been a vocal advocate for safeguarding authors’ works from unauthorised use in AI systems.

The ALCS recently conducted a survey of its members to gauge their views on AI, and the organisation welcomed PRH’s decision to address the issue explicitly in its copyright pages. ALCS CEO Barbara Hayes expressed encouragement that a major publisher like PRH is taking steps to reaffirm the principle of copyright and protect works from being used to train AI models.

“Major publishers adopting new wording in their printed materials to explicitly forbid the use of copyrighted works in AI training is a crucial step in safeguarding intellectual property,” according to Hayes. She stated that the ALCS hopes that more and more publishers will follow PRH’s lead and that technology companies will become aware of the new criteria.

However, while PRH’s efforts have been widely praised, others in the industry believe that more could be done. The Society of Authors (SoA) has welcomed PRH’s decision, but believes that simply revising copyright notices may not be enough.

X users sharing different perspectives in the ban on AI training
X users sharing different perspectives in the ban on AI training (Source – X)

The need for contractual protections

According to the SoA’s CEO, Anna Ganley, there is no standard wording for copyright protection, and while “All rights reserved” is a frequent phrase, it might be vague when used for more particular purposes, such as AI training. Ganley suggests that publishers need to go beyond copyright page disclaimers and include explicit language in author contracts that ensures authors’ rights are protected from AI exploitation.

Ganley explained, “We’re pleased to see publishers like PRH adding to the ‘All rights reserved’ notice to specifically exclude the use of a work for the purpose of training generative AI models. This provides greater clarity and helps readers understand what cannot be done without the rights-holder’s consent.” However, she stressed that contract language must also be revised to guarantee that authors are consulted before their works are used in AI-related projects, such as for narrating audiobooks, generating translations, or designing covers.

The broader concern reflects the complexities and evolving nature of AI’s impact on the creative industries. While book publishers are beginning to address the issues, AI scraping has also had an impact on journalism and the media. The New York Times filed a copyright infringement lawsuit against OpenAI and Microsoft in December 2023, saying their articles were used to train AI models without permission. According to the lawsuit, OpenAI’s models were capable of generating text that closely resembled the Times‘s original content, raising concerns about the legal boundaries of content use for AI training.

The Times claimed that while a wide variety of sources were used to train AI models, its content was given particular emphasis, reflecting the value that companies place on high-quality journalism. The case, along with PRH’s recent actions, emphasises the worries between content creators and AI developers, as well as the need for clearer legal frameworks in which to navigate challenges.

Looking ahead: AI and copyright conflicts

The conflict between AI and copyright is likely to intensify as AI systems grow more sophisticated and widespread. PRH’s decision to update its copyright notices may mark the beginning of a broader shift in how publishers and other content creators regard and interact with AI, but it also highlights the complexities of balancing technological innovation with intellectual property protection. It remains to be seen whether other major publishers will follow PRH’s lead, but the conversation over AI’s position in the creative industries is far from over.

Author

  • As a tech journalist, Zul focuses on topics including cloud computing, cybersecurity, and disruptive technology in the enterprise industry. He has expertise in moderating webinars and presenting content on video, in addition to having a background in networking technology.

    View all posts

About the Author

Muhammad Zulhusni

As a tech journalist, Zul focuses on topics including cloud computing, cybersecurity, and disruptive technology in the enterprise industry. He has expertise in moderating webinars and presenting content on video, in addition to having a background in networking technology.

Related

August 11, 2026

August 10, 2026

August 5, 2026

July 30, 2026

Join our Community

Subscribe now to get all our premium content and latest tech news delivered straight to your inbox

Popular

12345 view(s)
11326 view(s)
7643 view(s)
6152 view(s)

Subscribe

All our premium content and latest tech news delivered straight to your inbox

This field is for validation purposes and should be left unchanged.
Name(Required)