AI Safety & Guardrails

Course Content

AI Safety & Guardrails

5 sections · 50 lessons

How can AI systems avoid reproducing copyrighted content?


What you need to know

Why models reproduce text

Memorisation is when a model outputs a training sequence word for word. Research on large models found memorisation grows with model size, with the number of times a sequence appears in training data, and with how much of the prefix the prompt provides. Popular text — song lyrics, famous opening lines, common code — is duplicated many times on the web, so it is memorised most.

Training-side controls (if you train or fine-tune)

  • Deduplicate: exact and near-duplicate removal (MinHash, suffix arrays). The single highest-leverage fix.
  • Filter known protected corpora and honour opt-outs such as robots.txt rules for AI crawlers.
  • Measure memorisation: feed the first 50 tokens of known training documents and check whether the model continues them verbatim; track the longest exact span as a release metric.

Output-side controls (where most teams operate)

Python
def ngrams(text, n=8):    words = text.lower().split()    return {tuple(words[i:i + n]) for i in range(len(words) - n + 1)}def overlap(output, protected, n=8):    out = ngrams(output, n)    return len(out & ngrams(protected, n)) / max(len(out), 1)protected = ("it was the best of times it was the worst of times "             "it was the age of wisdom it was the age of foolishness")output = ("as the novel opens it was the best of times it was the worst "          "of times it was the age of wisdom")print(round(overlap(output, protected), 2))   # 0.73 -> block or rewrite

The check counts how many 8-word sequences in the output also appear in a protected text. Here 73% of them match, so the output is essentially a copy. In production you index the protected corpus with MinHash or a suffix array so you can check against millions of documents quickly, and set a threshold per use case.

Other output-side controls:

  • Cap quotation length in RAG and require attribution when quoting.
  • Retrieve only from licensed content, and keep licence metadata on every document so the answer can respect it.
  • Turn on vendor protections such as code-reference filters; indemnities usually depend on them.
  • Input constraints: refuse or rephrase prompts that ask for full lyrics, whole chapters, or "in the exact style of" a named living artist for commercial assets.

A real-life example

A news-summary app uses RAG over articles from 40 publishers, only 25 of which have signed licences. A publisher complains that the app reproduces entire paragraphs of its paywalled stories.

The team finds two causes: articles from unlicensed publishers were being crawled into the index, and the prompt said "use the article's wording where possible". The fixes: index only licensed publishers, tagged with licence terms; limit quotations to 40 words with a link to the source; run an 8-gram overlap check on every summary and regenerate if more than 20% of n-grams match a single source. On a 2,000-article test, summaries above the threshold fall from 14% to under 1%.

Follow-up questions to expect

  • "Does paraphrasing solve it?" — It reduces verbatim copying, but a close paraphrase of a whole creative work can still infringe. Summarising facts is safer than rewriting expression.
  • "How do you choose n?" — Short n-grams match common phrases and cause false alarms; 8 to 13 words is a common range for text. Tune it on examples your lawyers consider copying.
  • "What about code?" — Use the coding tool's public-code matching filter, and scan for licence headers and known snippets in CI.