Course Content
AI Safety & Guardrails
5 sections · 50 lessons
What are key copyright risks in AI-generated outputs?
What you need to know
1. Training data (mostly the provider's risk)
The legal position is still moving and depends on the country:
- US: courts are deciding fair use case by case. In 2025, rulings in cases against Anthropic and Meta found that training on books could be fair use on those facts, while a court also held that building a library from pirated copies was not protected; the Anthropic case then settled for a large sum.
- EU: text-and-data-mining is allowed unless rights-holders opt out in a machine-readable way, and the EU AI Act requires general-purpose model providers to have a copyright policy and publish a summary of training content. A German court ruled in late 2025 that a model which memorised and reproduced song lyrics infringed.
- India: no specific AI rule yet; a case by a news agency against OpenAI has been before the Delhi High Court.
Expect this to change; say "unsettled" in an interview rather than predicting outcomes.
2. Output infringement (your risk)
- Long verbatim passages from books, articles or lyrics.
- Recognisable characters or art styles close to a specific work.
- Code under a copyleft licence reproduced from training data, bringing licence obligations into your product. This is the version engineering teams hit most.
3. Ownership of outputs
The US Copyright Office's 2025 report said prompts alone do not usually give enough human control for authorship; human selection, arrangement or modification of the output can be protected. Courts upheld the refusal to register a work listing an AI as author. Other countries differ.
4. Adjacent rights
Trademarks (a generated logo that copies a brand) and personality rights — using a real person's voice, face or name. Indian courts have granted several celebrities orders protecting their personality rights against AI misuse.
Mitigations
- Vendor IP indemnities — read the conditions; they usually require you to keep the vendor's filters on and not intentionally prompt for infringing content.
- Output-side similarity checks against protected or licensed material.
- Records of human contribution for assets you need to own.
- Disclosure of AI use where contracts or regulators require it.
A real-life example
A marketing agency uses an image model to create a festive campaign for a snack brand. One generated mascot looks very close to a well-known cartoon character, because the brief asked for "a cheerful character in the style of" that cartoon. Legal stops it before launch.
The agency changes its process: prompts may not name artists, studios or characters; every final asset goes through a reverse image search; designers substantially redraw generated drafts and keep the layered source files as a record of human authorship; and the vendor contract's indemnity conditions are summarised on one page for the creative team.
Follow-up questions to expect
- "Who is liable if the model outputs infringing text?" — Usually the company that publishes or uses it; the vendor indemnity may shift some cost, only if its conditions were met.
- "Is 'in the style of' infringement?" — Style alone is generally not protected by copyright, but naming a specific work or character raises the chance of a substantially similar output, which can infringe.
- "How do you handle generated code licences?" — Turn on the tool's duplicate-code detection or reference filter, and scan with a licence-compliance tool as you would for copied open-source code.