|

Text And Data Mining Opt-Outs

Training a model on copyright-protected material can be entirely lawful in the EU. Whether it is depends on a single question: did the rightsholder reserve its rights. That single question decides the whole text and data mining analysis. AIGP scenarios put candidates on the wrong side of it constantly.

There are two text and data mining exceptions

The Copyright in the Digital Single Market Directive created two. They behave very differently, and mixing them up is the most common error on this subject.

Article 3 is the research exception. It permits reproductions and extractions by research organisations and cultural heritage institutions. The purpose must be scientific research. The works must be ones “to which they have lawful access”. Two limits sit on the face of it. Only those bodies benefit, and only scientific text and data mining qualifies.

Article 4 is the general exception. It covers “reproductions and extractions of lawfully accessible works and other subject matter for the purposes of text and data mining”. No restriction attaches to who mines, or why. Commercial use sits squarely in scope. A startup scraping a public web corpus exercises the same right as a university.

Then comes the difference that decides most exam items. Article 4(3) allows rightsholders to switch the general exception off. Nobody can opt out of research text and data mining. Article 3 carries no such power, and Article 4(4) says so expressly: the general exception “shall not affect the application of Article 3”.

What lawful access means

Both articles hang on access. Recital 14 explains the term. It covers content reached “based on an open access policy or through contractual arrangements between rightholders and research organisations or cultural heritage institutions, such as subscriptions, or through other lawful means”.

Access obtained by circumventing a paywall is not lawful access. Neither is access in breach of the terms you agreed to when you got it. The text and data mining exceptions do not repair a defective starting point.

How the text and data mining reservation works

Article 4(3) sets the condition precisely. The general exception applies on one condition. The use must not have “been expressly reserved by their rightholders in an appropriate manner, such as machine-readable means in the case of content made publicly available online”.

Read “appropriate manner” as the operative standard. Machine-readable means are an illustration of it, not the definition.

Online content is stricter than the article suggests

Recital 18 tightens this considerably. Take content made publicly available online. There, it “should only be considered appropriate to reserve those rights by the use of machine-readable means, including metadata and terms and conditions of a website or a service”.

A blog post declaring “no AI training here” in prose does nothing for text and data mining purposes. It has to sit in machine-readable form as well. Website terms and conditions do count, which surprises anyone assuming the reservation must live in a robots file.

The recital then adds the case everyone forgets. Where content has not been made publicly available online, other means work. The recital names “contractual agreements or a unilateral declaration”. Offline and licensed collections behave differently from the open web.

What the AI Act stacks on top

The AI Act does not amend copyright law. It converts the reservation into a compliance duty for one group.

Article 53(1)(c) puts a copyright policy on providers of general-purpose AI models. That policy must “identify and comply with, including through state-of-the-art technologies, a reservation of rights expressed pursuant to Article 4(3) of Directive (EU) 2019/790”. So the text and data mining reservation becomes an engineering requirement, not a legal footnote. Article 53(1)(d) then demands a sufficiently detailed public summary of the training content, on a template from the AI Office. The Commission published that template on 24 July 2025.

Recital 106 gives the duty its reach. Any provider placing a general-purpose model on the Union market carries it, whatever jurisdiction the training took place in. Mining performed lawfully somewhere else does not discharge the obligation.

These obligations have applied since 2 August 2025. Models already on the market at that date have until 2 August 2027 to comply. One currency point. The Digital Omnibus on AI amended the AI Act in July 2026 and moved several deadlines. It left Article 53 alone.

Where text and data mining law is still moving

Three things remain genuinely open, and AIGP candidates should know the shape of each rather than the outcome.

The Commission consulted on protocols for reserving rights from text and data mining between December 2025 and January 2026. It said it would publish a list of generally agreed machine-readable opt-out solutions. No such list has appeared. Until one does, “machine-readable” remains a standard without an agreed dictionary.

No CJEU judgment interprets Article 4. A reference is pending in Case C-250/25, Like Company v Google Ireland. It asks the Court directly whether reproduction for model training falls inside the exception.

The German litigation is further along and worth tracking. The Hamburg Regional Court dismissed the LAION claim on 27 September 2024. Appeal to the Hanseatic Higher Regional Court failed on 10 December 2025. It now sits with the Federal Court of Justice as I ZR 281/25, with a hearing listed for 3 September 2026. No German court has ruled finally on whether the general exception covers dataset creation for generative models.

How this lands in the exam

The Body of Knowledge is the IAPP’s published map of what each exam covers. Text and data mining sits in the domain on existing law that reaches AI. Questions rarely ask what Article 4 says. They describe a sourcing decision and ask what makes it defensible.

Three habits will carry you through most of them. Check access before purpose, because unlawful access defeats both exceptions. Then ask who is mining, since the research exception is closed to commercial developers. Last, look for a reservation in machine-readable form, because a prose notice on a web page fails the online test.

Sourcing defects also travel. A model trained on improperly acquired material carries that exposure into deployment. The same logic sits behind product liability for AI systems. It also makes the choice between build, buy and adapt a governance decision rather than a procurement one.

Fifteen minutes on the free AIGP assessment will tell you whether the legal domain is where your marks are leaking. If the questions themselves keep outmanoeuvring you, the AIGP Exam Question Masterclass works on that directly.

Similar Posts