Course Content
AI Safety & Guardrails
5 sections · 50 lessons
How do knowledge graphs improve grounding beyond standard retrieval?
What you need to know
What a graph gives you that chunks don't
- Multi-hop: "Which of our suppliers are owned by a company sanctioned last year?" is two joins. Chunk retrieval finds fragments and the model guesses the link.
- Completeness: "List all contracts expiring this quarter" needs the full set; top-k returns only k.
- Constraints: dates, versions and hierarchy are fields you filter on, not words the embedding may ignore.
- Provenance per fact: each edge records its source document, so citations are exact.
- Contradictions: a rule that a person has one date of birth makes conflicts visible.
- Permissions at node and edge level.
Two common patterns
- Text-to-query: the LLM writes a Cypher or SPARQL query from the question; the database runs it; the model answers from the returned rows. Validate generated queries (read-only, allowed labels, row limits).
- GraphRAG-style (popularised by Microsoft's GraphRAG): extract entities and relations from documents, build community summaries, and retrieve over both the graph and the text.
Costs
- Extraction errors: an LLM that extracts relations makes mistakes, and the graph then presents them as structured facts.
- Schema design and maintenance is ongoing work.
- Staleness looks authoritative: an outdated edge is still a clean, confident fact.
A real-life example
A healthcare symptom-checker is asked, "I take warfarin — can I take ibuprofen for this headache?" Chunk retrieval returns a general ibuprofen leaflet and a general warfarin leaflet; the model sometimes misses the interaction.
The team adds a drug-interaction graph built from a licensed clinical database: (warfarin)-[INTERACTS_WITH {severity: "major", source: "…"}]->(ibuprofen). For any question naming two or more drugs, a deterministic query checks all pairs, and a major interaction always produces a warning plus "talk to your doctor", with the source cited. On 300 interaction test cases, missed major interactions fall from 12% to 0%.
Follow-up questions to expect
- "When is a graph not worth it?" — When questions are mostly open-ended and single-document ("summarise this policy"). The build and maintenance cost only pays off for structured, relational questions.
- "How do you stop text-to-Cypher from doing damage?" — Run it with a read-only user, restrict labels and depth, add a row limit and timeout, and validate the query before execution.
- "How do you keep the graph fresh?" — Incremental extraction on document change, timestamps on edges, and scheduled checks for edges whose source was deleted.