Course Content
AI Agent Frameworks
4 sections · 15 lessons
Memory-Driven Workflows with Semantic Kernel
A support assistant was built over 340 internal help articles. Each article — some of them twelve pages long — was stored as one memory record. Retrieval returned the top three articles.
A customer asked: "How long do I have to return a faulty item bought on finance?" The assistant returned a confident, wrong answer: 14 days. The correct answer, 30 days for finance purchases, was in the retrieved article — on page nine, in a subsection about credit agreements.
Two things had gone wrong at once. The three retrieved articles totalled roughly 24,000 tokens, so the relevant sentence was one line in a wall of text. And the embedding for a twelve-page article is an average of everything in it: returns policy, delivery, warranty, credit terms, complaints procedure. Averaged together, that vector is close to nothing in particular. The specific question about finance returns matched it only weakly, and matched a shorter, more focused article about standard returns much more strongly.
An embedding of a long document is a blurred average of everything in it. The fix is never a better model — it is smaller units.
Memory in Semantic Kernel is straightforward to switch on and easy to get quietly wrong in exactly this way. What follows is how the machinery works, and where the accuracy actually comes from.
Two kinds of memory, and picking the wrong one costs you
| Semantic memory | Structured memory | |
|---|---|---|
| Stored as | Text plus an embedding vector | Rows, columns, keys |
| Queried by | Meaning: "what did we say about refunds?" | Exact match: WHERE order_id = 4471 |
| Returns | The k most similar items with scores | Exactly the matching rows |
| Correct answer guaranteed | No — similarity is approximate | Yes |
| Right for | Documents, past conversations, notes, policies | Order status, balances, entitlements, prices |
The failure mode of confusing them is severe. Ask a vector store "what is the balance on account 88213?" and it returns the three account records whose text is most similar to your query — which may well be accounts 88214 and 88231, because their text is nearly identical. It will return them with high confidence scores and no indication that they are the wrong accounts.
The rule: anything with an identifier goes in structured storage. Semantic search is for text where you do not know the exact words you are looking for.
Wiring memory into the kernel
pip install semantic-kernel1import os2from semantic_kernel import Kernel3from semantic_kernel.connectors.ai.open_ai import (4 OpenAIChatCompletion, OpenAITextEmbedding)5from semantic_kernel.memory import SemanticTextMemory, VolatileMemoryStore6from semantic_kernel.core_plugins import TextMemoryPlugin78kernel = Kernel()9kernel.add_service(OpenAIChatCompletion(10 service_id="chat", ai_model_id=os.getenv("CHAT_MODEL", "gpt-6-sol"),11 api_key=KEY))1213embedder = OpenAITextEmbedding(14 service_id="embed", ai_model_id="text-embedding-3-small", api_key=KEY)15kernel.add_service(embedder)1617memory = SemanticTextMemory(storage=VolatileMemoryStore(),18 embeddings_generator=embedder)19kernel.add_plugin(TextMemoryPlugin(memory), plugin_name="memory")One warning before anything else: SemanticTextMemory, VolatileMemoryStore and TextMemoryPlugin are Semantic Kernel's original memory API, and current releases mark them as deprecated and due for removal. They are used here because they show the mechanics in the fewest lines; the typed vector store API at the end of this lesson is the one to build on.
VolatileMemoryStore is an in-process dictionary. It vanishes on restart and is invisible to any other instance of your service. It is correct for development and for tests, and wrong for anything else — the moment you run two replicas, half your users get an assistant with no memory.
TextMemoryPlugin registers memory.recall and memory.save as callable functions, which means the model can decide to search memory as part of automatic function calling. That is convenient and it is also a decision worth making deliberately: a model that can call recall will sometimes not bother, and for a policy assistant you usually want retrieval to happen unconditionally before the model sees the question at all.
Saving and searching
1await memory.save_information(2 collection="policies",3 id="returns-finance-01",4 text=("Items bought on a finance agreement may be returned within 30 days "5 "of delivery if faulty. The finance agreement is cancelled from the "6 "date of return and no further instalments are due."),7 description="Returns window for finance purchases",8 additional_metadata='{"doc": "returns-policy", "page": 9, "version": "2025-04"}',9)1011hits = await memory.search(collection="policies",12 query="how long to return a faulty item on finance",13 limit=3, min_relevance_score=0.75)14for h in hits:15 print(f"{h.relevance:.3f} {h.id} {h.text[:80]}")min_relevance_score is the parameter people leave at its default and later regret. Without it, a query with no good match returns the three least-bad records, and the model treats them as evidence. With a floor of 0.75, a query about something you have never stored returns nothing, and the assistant can honestly say it does not know.
What the score means
Relevance is cosine similarity: the cosine of the angle between the query vector and the record vector. For vectors a and b:
A toy two-dimensional example makes the behaviour concrete. Take a=[0.6, 0.8] and b=[0.8, 0.6]. Both have length 0.36+0.64=1. The dot product is 0.6×0.8+0.8×0.6=0.48+0.48=0.96, so the similarity is 0.96 — very close, though the vectors are not identical.
Now c=[0.8, −0.6]: the dot product with a is 0.48−0.48=0, similarity 0 — unrelated. Real embeddings have hundreds or thousands of dimensions rather than two — text-embedding-3-small produces 1,536 — but the arithmetic and the interpretation are identical.
| Score band | Typical meaning | Suggested handling |
|---|---|---|
| 0.90+ | Near-paraphrase | Use directly |
| 0.80–0.90 | Same topic, possibly different aspect | Use, but let the model check relevance |
| 0.70–0.80 | Loosely related | Include only if nothing better |
| Below 0.70 | Usually noise | Discard; say you do not know |
These bands shift with the embedding model, so calibrate rather than copy: store twenty known records, run twenty queries whose correct answers you know, and read the score distribution of the correct hits versus the incorrect ones. The threshold sits between the two clusters.
Chunking — where the accuracy actually comes from
Back to the opening failure. A twelve-page article is about 6,000 words, roughly 8,000 tokens. Stored whole, it is one vector and one retrieval unit.
1def chunk(text: str, size: int = 400, overlap: int = 80) -> list[str]:2 """Split into ~`size`-token chunks with `overlap` tokens of context carried3 across each boundary. Sizes are in words as a rough token proxy."""4 words = text.split()5 step = size - overlap6 out = []7 for start in range(0, max(len(words) - overlap, 1), step):8 piece = words[start:start + size]9 if len(piece) < 40 and out: # merge a tiny tail into the last10 out[-1] = out[-1] + " " + " ".join(piece)11 break12 out.append(" ".join(piece))13 return outWork through the numbers for the 8,000-token article with size=400, overlap=80, so step=320:
| Whole document | Chunked | |
|---|---|---|
| Records stored | 1 | 25 |
| Tokens per record | 8,000 | 400 |
| Tokens injected for top-3 | 24,000 | 1,200 |
| Embedding covers | Everything, averaged | One topic |
| Cost per query at 3 dollars/million | 0.072 dollars | 0.0036 dollars |
The chunk count comes from ⌈(8000−400)/320⌉+1=⌈23.75⌉+1=25. A twentyfold reduction in injected context, and — more importantly — the retrieved 1,200 tokens are about finance returns rather than about everything the article covers.
Why overlap is not optional
Split on an exact boundary and a sentence spanning it is destroyed:
chunk 7 ends: "...may be returned within 30 days of delivery if"chunk 8 begins: "faulty. The finance agreement is cancelled from..."Neither chunk answers the question. With 80 tokens of overlap the sentence appears whole in chunk 8, and the cost is 25 per cent more storage — an entirely reasonable price. The rule of thumb: overlap of 15–25 per cent of chunk size.
Chunk size is a trade-off, not a constant
| Chunk size | Retrieval precision | Context completeness | Suits |
|---|---|---|---|
| 100–200 tokens | High | Poor — answers get truncated | FAQ pairs, definitions |
| 300–500 tokens | Good | Good | Policy documents, articles |
| 800–1,200 tokens | Falls — averaging returns | Excellent | Narrative, code files |
Where the document has real structure, split on it rather than by word count. A policy document with numbered clauses should chunk at clause boundaries; each clause is a self-contained unit and its embedding is naturally about one thing.
A memory-driven workflow
1from semantic_kernel.functions import KernelArguments23class SupportAssistant:4 def __init__(self, kernel, memory):5 self.kernel, self.memory = kernel, memory67 async def ingest(self, doc_id: str, title: str, body: str, version: str):8 for i, piece in enumerate(chunk(body, size=400, overlap=80)):9 await self.memory.save_information(10 collection="policies",11 id=f"{doc_id}#{i:03d}",12 text=f"[{title}] {piece}", # title in every chunk13 description=title,14 additional_metadata=f'{{"doc":"{doc_id}","chunk":{i},'15 f'"version":"{version}"}}',16 )1718 async def answer(self, question: str) -> str:19 hits = await self.memory.search(20 collection="policies", query=question,21 limit=4, min_relevance_score=0.78)2223 if not hits:24 return ("I could not find anything in the policy documents about "25 "that. Please ask a colleague to check.")2627 context = "\n\n".join(28 f"[{h.id} score={h.relevance:.2f}]\n{h.text}" for h in hits)2930 answer = await self.kernel.invoke(31 plugin_name="support", function_name="answer_from_policy",32 arguments=KernelArguments(context=context, question=question))33 return str(answer)Two details in ingest earn their place. Prefixing every chunk with the document title means a chunk retrieved out of context still says what it belongs to — without it, a retrieved paragraph about "30 days" gives the model no clue whether it concerns returns, refunds or cancellations. And storing the version in metadata is what lets you delete an entire superseded document later; without a version key, updating a policy means finding every chunk of it by hand.
The prompt function does the rest of the work:
1template = (2 "Answer using ONLY the policy extracts below.\n"3 "Rules:\n"4 "- If the extracts do not contain the answer, say so. Do not infer.\n"5 "- Quote the exact sentence you relied on, then explain it.\n"6 "- Cite the extract id in square brackets.\n"7 "- If two extracts conflict, say so and quote both.\n\n"8 "Extracts:\n{{$context}}\n\nQuestion: {{$question}}"9)"Quote the exact sentence you relied on" is the highest-value line. It makes hallucination visible: if the quoted sentence is not in the extracts, you can detect that programmatically with a substring check, and refuse the answer.
Two patterns worth knowing
Memory as a checkpoint between stages
A long pipeline — research, analyse, draft, review — where each stage writes its output to memory rather than passing it directly. Stage 3 can then retrieve stage 1's findings even though stage 2 sat between them, and a crash after stage 3 loses only stage 4.
1async def stage(name, fn, run_id, memory, query=None):2 prior = ""3 if query:4 hits = await memory.search(collection=f"run-{run_id}",5 query=query, limit=3)6 prior = "\n".join(h.text for h in hits)7 out = await fn(prior)8 await memory.save_information(collection=f"run-{run_id}",9 id=f"{name}", text=out, description=name)10 return outThe cost saving is real: the draft stage retrieves the 600 tokens it needs rather than carrying 4,000 tokens of accumulated pipeline output.
Memory-augmented decisions
Before acting, check what happened last time. A deployment agent that searches memory for "deploy failures for service X" and finds "last two deploys of X failed on migration timeout" behaves differently from one that starts fresh — and the difference costs one embedding call, about 0.00002 dollars.
1async def decide(action: str, memory) -> str:2 past = await memory.search(collection="outcomes",3 query=f"outcome of {action}",4 limit=3, min_relevance_score=0.80)5 if not past:6 return "no prior experience"7 failures = [p for p in past if "FAILED" in p.text]8 if len(failures) >= 2:9 return f"CAUTION: {len(failures)} prior failures. " + failures[0].text10 return "prior attempts succeeded"The vector store API
The SemanticTextMemory interface above is the older, simpler surface, and it is deprecated in favour of a typed vector-store model that supports real backends — Azure AI Search, Postgres with pgvector, Redis, Qdrant, Chroma — behind one interface. The same connectors also work from Microsoft Agent Framework, so this is the part of the lesson that carries forward.
1from dataclasses import dataclass2from typing import Annotated3from semantic_kernel.data.vector import VectorStoreField, vectorstoremodel45@vectorstoremodel6@dataclass7class PolicyChunk:8 id: Annotated[str, VectorStoreField("key")]9 text: Annotated[str, VectorStoreField("data", is_full_text_indexed=True)]10 doc_id: Annotated[str, VectorStoreField("data", is_indexed=True)]11 version: Annotated[str, VectorStoreField("data", is_indexed=True)]12 vector: Annotated[list[float] | None,13 VectorStoreField("vector", dimensions=1536)] = Noneis_indexed=True on doc_id and version is the capability worth migrating for. It lets you combine semantic search with exact filters — "find chunks similar to this question, but only from version 2025-04" — which the older interface cannot express. Without it, a superseded policy competes with the current one on similarity alone, and the model has no way to tell which is in force.
Semantic search returns the closest thing it has, never nothing. Without a relevance floor, "I do not know" is an answer your assistant is structurally incapable of giving.
Failure modes with names
| Symptom | Cause | Fix |
|---|---|---|
| Confident answer from the wrong document | No min_relevance_score; the least-bad match was returned | Set a floor; return "not found" below it |
| Right document, wrong detail | Chunks too large; embedding is an average | Chunk to 300–500 tokens on structural boundaries |
| Answer cut off mid-rule | No overlap; the sentence spans a boundary | 15–25 per cent overlap |
| Retrieved text lacks context | Chunk stored without its title or section | Prefix every chunk with document and section |
| Old policy quoted after an update | Superseded chunks never deleted | Version metadata plus filtered search |
| Memory empty after restart | VolatileMemoryStore (or any in-memory store) in production | A persistent connector |
| Wrong customer's data returned | Semantic search used for identifier lookup | Structured store for anything with an id |
| Retrieval quality falls as the corpus grows | More near-duplicates competing at the same score | Deduplicate on ingest; add filters to narrow the candidate set |
Building this so it stays correct
Memory systems degrade silently. There is no exception when retrieval returns the wrong chunk — just an answer that is slightly, confidently wrong, which nobody notices until a customer acts on it. Three practices catch that.
Keep a retrieval test set. Thirty questions, each with the id of the chunk that should be retrieved. Run it after every ingest, every chunk-size change and every embedding-model change. Report recall at 3 — how often the correct chunk appears in the top three. If it drops from 27/30 to 22/30 after you ingested a new document set, you have found a problem before your users did.
Log the scores of what you retrieve, always. When someone reports a wrong answer, the first question is what was retrieved and at what score. If the correct chunk scored 0.74 and your floor was 0.78, the fix is the floor. If the correct chunk was not in the store at all, the fix is ingestion. These are entirely different problems and you cannot tell them apart without the log.
Make the answer quote its source. Requiring an exact quotation from the retrieved extracts turns hallucination from an invisible failure into a checkable one — a substring test on the quote against the context is three lines of code and catches the class of error that does the most damage. An assistant that says "I could not find this in the policy documents" is doing its job. One that invents 14 days is not, and the only difference between them is whether you built the check.