AI Agent Frameworks

Memory-Driven Workflows with Semantic Kernel


A support assistant was built over 340 internal help articles. Each article — some of them twelve pages long — was stored as one memory record. Retrieval returned the top three articles.

A customer asked: "How long do I have to return a faulty item bought on finance?" The assistant returned a confident, wrong answer: 14 days. The correct answer, 30 days for finance purchases, was in the retrieved article — on page nine, in a subsection about credit agreements.

Two things had gone wrong at once. The three retrieved articles totalled roughly 24,000 tokens, so the relevant sentence was one line in a wall of text. And the embedding for a twelve-page article is an average of everything in it: returns policy, delivery, warranty, credit terms, complaints procedure. Averaged together, that vector is close to nothing in particular. The specific question about finance returns matched it only weakly, and matched a shorter, more focused article about standard returns much more strongly.

An embedding of a long document is a blurred average of everything in it. The fix is never a better model — it is smaller units.

Memory in Semantic Kernel is straightforward to switch on and easy to get quietly wrong in exactly this way. What follows is how the machinery works, and where the accuracy actually comes from.

Chunking a twelve-page article12345678910111201234567891011chunkstartsoverlapOverlap keeps a passage whole when the answer straddles a boundary.
One record per article means every search scores twelve pages at once — the relevant paragraph is averaged away.

Two kinds of memory, and picking the wrong one costs you

Semantic memoryStructured memory
Stored asText plus an embedding vectorRows, columns, keys
Queried byMeaning: "what did we say about refunds?"Exact match: WHERE order_id = 4471
ReturnsThe k most similar items with scoresExactly the matching rows
Correct answer guaranteedNo — similarity is approximateYes
Right forDocuments, past conversations, notes, policiesOrder status, balances, entitlements, prices

The failure mode of confusing them is severe. Ask a vector store "what is the balance on account 88213?" and it returns the three account records whose text is most similar to your query — which may well be accounts 88214 and 88231, because their text is nearly identical. It will return them with high confidence scores and no indication that they are the wrong accounts.

The rule: anything with an identifier goes in structured storage. Semantic search is for text where you do not know the exact words you are looking for.

Wiring memory into the kernel

Bash
pip install semantic-kernel
Python
import osfrom semantic_kernel import Kernelfrom semantic_kernel.connectors.ai.open_ai import (    OpenAIChatCompletion, OpenAITextEmbedding)from semantic_kernel.memory import SemanticTextMemory, VolatileMemoryStorefrom semantic_kernel.core_plugins import TextMemoryPluginkernel = Kernel()kernel.add_service(OpenAIChatCompletion(    service_id="chat", ai_model_id=os.getenv("CHAT_MODEL", "gpt-6-sol"),    api_key=KEY))embedder = OpenAITextEmbedding(    service_id="embed", ai_model_id="text-embedding-3-small", api_key=KEY)kernel.add_service(embedder)memory = SemanticTextMemory(storage=VolatileMemoryStore(),                            embeddings_generator=embedder)kernel.add_plugin(TextMemoryPlugin(memory), plugin_name="memory")

One warning before anything else: SemanticTextMemory, VolatileMemoryStore and TextMemoryPlugin are Semantic Kernel's original memory API, and current releases mark them as deprecated and due for removal. They are used here because they show the mechanics in the fewest lines; the typed vector store API at the end of this lesson is the one to build on.

VolatileMemoryStore is an in-process dictionary. It vanishes on restart and is invisible to any other instance of your service. It is correct for development and for tests, and wrong for anything else — the moment you run two replicas, half your users get an assistant with no memory.

TextMemoryPlugin registers memory.recall and memory.save as callable functions, which means the model can decide to search memory as part of automatic function calling. That is convenient and it is also a decision worth making deliberately: a model that can call recall will sometimes not bother, and for a policy assistant you usually want retrieval to happen unconditionally before the model sees the question at all.

Saving and searching

Python
await memory.save_information(    collection="policies",    id="returns-finance-01",    text=("Items bought on a finance agreement may be returned within 30 days "          "of delivery if faulty. The finance agreement is cancelled from the "          "date of return and no further instalments are due."),    description="Returns window for finance purchases",    additional_metadata='{"doc": "returns-policy", "page": 9, "version": "2025-04"}',)hits = await memory.search(collection="policies",                           query="how long to return a faulty item on finance",                           limit=3, min_relevance_score=0.75)for h in hits:    print(f"{h.relevance:.3f}  {h.id}  {h.text[:80]}")

min_relevance_score is the parameter people leave at its default and later regret. Without it, a query with no good match returns the three least-bad records, and the model treats them as evidence. With a floor of 0.75, a query about something you have never stored returns nothing, and the assistant can honestly say it does not know.

What the score means

Relevance is cosine similarity: the cosine of the angle between the query vector and the record vector. For vectors a\mathbf{a} and b\mathbf{b}:

cos⁡θ=a⋅b∥a∥ ∥b∥\cos\theta = \frac{\mathbf{a} \cdot \mathbf{b}}{\lVert\mathbf{a}\rVert \, \lVert\mathbf{b}\rVert}

A toy two-dimensional example makes the behaviour concrete. Take a=[0.6, 0.8]\mathbf{a} = [0.6,\ 0.8] and b=[0.8, 0.6]\mathbf{b} = [0.8,\ 0.6]. Both have length 0.36+0.64=1\sqrt{0.36 + 0.64} = 1. The dot product is 0.6×0.8+0.8×0.6=0.48+0.48=0.960.6 \times 0.8 + 0.8 \times 0.6 = 0.48 + 0.48 = 0.96, so the similarity is 0.96 — very close, though the vectors are not identical.

Now c=[0.8, −0.6]\mathbf{c} = [0.8,\ -0.6]: the dot product with a\mathbf{a} is 0.48−0.48=00.48 - 0.48 = 0, similarity 0 — unrelated. Real embeddings have hundreds or thousands of dimensions rather than two — text-embedding-3-small produces 1,536 — but the arithmetic and the interpretation are identical.

Score bandTypical meaningSuggested handling
0.90+Near-paraphraseUse directly
0.80–0.90Same topic, possibly different aspectUse, but let the model check relevance
0.70–0.80Loosely relatedInclude only if nothing better
Below 0.70Usually noiseDiscard; say you do not know

These bands shift with the embedding model, so calibrate rather than copy: store twenty known records, run twenty queries whose correct answers you know, and read the score distribution of the correct hits versus the incorrect ones. The threshold sits between the two clusters.

Chunking — where the accuracy actually comes from

Back to the opening failure. A twelve-page article is about 6,000 words, roughly 8,000 tokens. Stored whole, it is one vector and one retrieval unit.

Python
def chunk(text: str, size: int = 400, overlap: int = 80) -> list[str]:    """Split into ~`size`-token chunks with `overlap` tokens of context carried    across each boundary. Sizes are in words as a rough token proxy."""    words = text.split()    step = size - overlap    out = []    for start in range(0, max(len(words) - overlap, 1), step):        piece = words[start:start + size]        if len(piece) < 40 and out:          # merge a tiny tail into the last            out[-1] = out[-1] + " " + " ".join(piece)            break        out.append(" ".join(piece))    return out

Work through the numbers for the 8,000-token article with size=400, overlap=80, so step=320:

Whole documentChunked
Records stored125
Tokens per record8,000400
Tokens injected for top-324,0001,200
Embedding coversEverything, averagedOne topic
Cost per query at 3 dollars/million0.072 dollars0.0036 dollars

The chunk count comes from ⌈(8000−400)/320⌉+1=⌈23.75⌉+1=25\lceil (8000 - 400)/320 \rceil + 1 = \lceil 23.75 \rceil + 1 = 25. A twentyfold reduction in injected context, and — more importantly — the retrieved 1,200 tokens are about finance returns rather than about everything the article covers.

Why overlap is not optional

Split on an exact boundary and a sentence spanning it is destroyed:

Text
chunk 7 ends:   "...may be returned within 30 days of delivery if"chunk 8 begins: "faulty. The finance agreement is cancelled from..."

Neither chunk answers the question. With 80 tokens of overlap the sentence appears whole in chunk 8, and the cost is 25 per cent more storage — an entirely reasonable price. The rule of thumb: overlap of 15–25 per cent of chunk size.

Chunk size is a trade-off, not a constant

Chunk sizeRetrieval precisionContext completenessSuits
100–200 tokensHighPoor — answers get truncatedFAQ pairs, definitions
300–500 tokensGoodGoodPolicy documents, articles
800–1,200 tokensFalls — averaging returnsExcellentNarrative, code files

Where the document has real structure, split on it rather than by word count. A policy document with numbered clauses should chunk at clause boundaries; each clause is a self-contained unit and its embedding is naturally about one thing.

A memory-driven workflow

Python
from semantic_kernel.functions import KernelArgumentsclass SupportAssistant:    def __init__(self, kernel, memory):        self.kernel, self.memory = kernel, memory    async def ingest(self, doc_id: str, title: str, body: str, version: str):        for i, piece in enumerate(chunk(body, size=400, overlap=80)):            await self.memory.save_information(                collection="policies",                id=f"{doc_id}#{i:03d}",                text=f"[{title}] {piece}",          # title in every chunk                description=title,                additional_metadata=f'{{"doc":"{doc_id}","chunk":{i},'                                    f'"version":"{version}"}}',            )    async def answer(self, question: str) -> str:        hits = await self.memory.search(            collection="policies", query=question,            limit=4, min_relevance_score=0.78)        if not hits:            return ("I could not find anything in the policy documents about "                    "that. Please ask a colleague to check.")        context = "\n\n".join(            f"[{h.id} score={h.relevance:.2f}]\n{h.text}" for h in hits)        answer = await self.kernel.invoke(            plugin_name="support", function_name="answer_from_policy",            arguments=KernelArguments(context=context, question=question))        return str(answer)

Two details in ingest earn their place. Prefixing every chunk with the document title means a chunk retrieved out of context still says what it belongs to — without it, a retrieved paragraph about "30 days" gives the model no clue whether it concerns returns, refunds or cancellations. And storing the version in metadata is what lets you delete an entire superseded document later; without a version key, updating a policy means finding every chunk of it by hand.

The prompt function does the rest of the work:

Python
template = (    "Answer using ONLY the policy extracts below.\n"    "Rules:\n"    "- If the extracts do not contain the answer, say so. Do not infer.\n"    "- Quote the exact sentence you relied on, then explain it.\n"    "- Cite the extract id in square brackets.\n"    "- If two extracts conflict, say so and quote both.\n\n"    "Extracts:\n{{$context}}\n\nQuestion: {{$question}}")

"Quote the exact sentence you relied on" is the highest-value line. It makes hallucination visible: if the quoted sentence is not in the extracts, you can detect that programmatically with a substring check, and refuse the answer.

Two patterns worth knowing

Memory as a checkpoint between stages

A long pipeline — research, analyse, draft, review — where each stage writes its output to memory rather than passing it directly. Stage 3 can then retrieve stage 1's findings even though stage 2 sat between them, and a crash after stage 3 loses only stage 4.

Python
async def stage(name, fn, run_id, memory, query=None):    prior = ""    if query:        hits = await memory.search(collection=f"run-{run_id}",                                   query=query, limit=3)        prior = "\n".join(h.text for h in hits)    out = await fn(prior)    await memory.save_information(collection=f"run-{run_id}",                                  id=f"{name}", text=out, description=name)    return out

The cost saving is real: the draft stage retrieves the 600 tokens it needs rather than carrying 4,000 tokens of accumulated pipeline output.

Memory-augmented decisions

Before acting, check what happened last time. A deployment agent that searches memory for "deploy failures for service X" and finds "last two deploys of X failed on migration timeout" behaves differently from one that starts fresh — and the difference costs one embedding call, about 0.00002 dollars.

Python
async def decide(action: str, memory) -> str:    past = await memory.search(collection="outcomes",                               query=f"outcome of {action}",                               limit=3, min_relevance_score=0.80)    if not past:        return "no prior experience"    failures = [p for p in past if "FAILED" in p.text]    if len(failures) >= 2:        return f"CAUTION: {len(failures)} prior failures. " + failures[0].text    return "prior attempts succeeded"

The vector store API

The SemanticTextMemory interface above is the older, simpler surface, and it is deprecated in favour of a typed vector-store model that supports real backends — Azure AI Search, Postgres with pgvector, Redis, Qdrant, Chroma — behind one interface. The same connectors also work from Microsoft Agent Framework, so this is the part of the lesson that carries forward.

Python
from dataclasses import dataclassfrom typing import Annotatedfrom semantic_kernel.data.vector import VectorStoreField, vectorstoremodel@vectorstoremodel@dataclassclass PolicyChunk:    id: Annotated[str, VectorStoreField("key")]    text: Annotated[str, VectorStoreField("data", is_full_text_indexed=True)]    doc_id: Annotated[str, VectorStoreField("data", is_indexed=True)]    version: Annotated[str, VectorStoreField("data", is_indexed=True)]    vector: Annotated[list[float] | None,                      VectorStoreField("vector", dimensions=1536)] = None

is_indexed=True on doc_id and version is the capability worth migrating for. It lets you combine semantic search with exact filters — "find chunks similar to this question, but only from version 2025-04" — which the older interface cannot express. Without it, a superseded policy competes with the current one on similarity alone, and the model has no way to tell which is in force.

Semantic search returns the closest thing it has, never nothing. Without a relevance floor, "I do not know" is an answer your assistant is structurally incapable of giving.

Failure modes with names

SymptomCauseFix
Confident answer from the wrong documentNo min_relevance_score; the least-bad match was returnedSet a floor; return "not found" below it
Right document, wrong detailChunks too large; embedding is an averageChunk to 300–500 tokens on structural boundaries
Answer cut off mid-ruleNo overlap; the sentence spans a boundary15–25 per cent overlap
Retrieved text lacks contextChunk stored without its title or sectionPrefix every chunk with document and section
Old policy quoted after an updateSuperseded chunks never deletedVersion metadata plus filtered search
Memory empty after restartVolatileMemoryStore (or any in-memory store) in productionA persistent connector
Wrong customer's data returnedSemantic search used for identifier lookupStructured store for anything with an id
Retrieval quality falls as the corpus growsMore near-duplicates competing at the same scoreDeduplicate on ingest; add filters to narrow the candidate set

Building this so it stays correct

Memory systems degrade silently. There is no exception when retrieval returns the wrong chunk — just an answer that is slightly, confidently wrong, which nobody notices until a customer acts on it. Three practices catch that.

Keep a retrieval test set. Thirty questions, each with the id of the chunk that should be retrieved. Run it after every ingest, every chunk-size change and every embedding-model change. Report recall at 3 — how often the correct chunk appears in the top three. If it drops from 27/30 to 22/30 after you ingested a new document set, you have found a problem before your users did.

Log the scores of what you retrieve, always. When someone reports a wrong answer, the first question is what was retrieved and at what score. If the correct chunk scored 0.74 and your floor was 0.78, the fix is the floor. If the correct chunk was not in the store at all, the fix is ingestion. These are entirely different problems and you cannot tell them apart without the log.

Make the answer quote its source. Requiring an exact quotation from the retrieved extracts turns hallucination from an invisible failure into a checkable one — a substring test on the quote against the context is three lines of code and catches the class of error that does the most damage. An assistant that says "I could not find this in the policy documents" is doing its job. One that invents 14 days is not, and the only difference between them is whether you built the check.