Course Content
AI Agent Frameworks
4 sections · 15 lessons
Semantic Kernel Foundations — Kernel, Plugins, and Planners
An insurance company has 340 internal services. Policy lookup, claims history, premium calculation, document generation, fraud scoring — all of them already exist, all of them audited, all of them with an owner who will not let you reimplement their logic in a prompt.
They want an assistant that answers "Has policy 88213 had any claims in the last two years, and what would the renewal premium be at the new rate?" Answering that means calling three existing services in the right order and doing one calculation.
The naive approach is to describe those services to a model and hope. The problem is not that it will not work — it will, roughly. The problem is that the compliance team has questions: which functions is the assistant allowed to call, who approved that list, what arguments were passed on 14 March, and can we turn off the premium calculator for a week without redeploying?
Semantic Kernel is Microsoft's answer, and its shape follows directly from those questions. Where other frameworks treat tools as a list you hand to an agent, Semantic Kernel treats them as a registry owned by a kernel object: functions are registered into named plugins, the registry is inspectable, and the same registry serves both automatic model-driven calling and direct programmatic invocation.
Most agent frameworks give the model a list of tools. Semantic Kernel gives your application a registry of functions, and lets the model call into it. That inversion is the whole design.
The kernel object
The kernel is a container holding three kinds of thing.
| Contents | What it is | Example |
|---|---|---|
| AI services | Chat, embedding or image models, each with a service_id | A chat model, a cheaper chat model, an embedder |
| Plugins | Named collections of callable functions | policy, claims, maths |
| Filters and settings | Cross-cutting hooks and execution config | Logging, auth checks, retry policy |
pip install semantic-kernel1import asyncio, os2from semantic_kernel import Kernel3from semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletion45kernel = Kernel()67kernel.add_service(OpenAIChatCompletion(8 service_id="fast",9 ai_model_id=os.getenv("FAST_MODEL", "gpt-6-luna"),10 api_key=API_KEY,11))12kernel.add_service(OpenAIChatCompletion(13 service_id="smart",14 ai_model_id=os.getenv("SMART_MODEL", "gpt-6-sol"),15 api_key=API_KEY,16))Two services registered, addressable by id. That is not a convenience feature — it is how you route classification to a model costing a fraction of the price and reserve the expensive one for reasoning. Registering both at kernel level means the choice is made per function, in configuration, rather than baked into code.
Functions: the unit everything is built from
Semantic Kernel has exactly one abstraction for "a thing that can be called", and both flavours share it.
| Native function | Prompt function | |
|---|---|---|
| Body | Python code | A prompt template |
| Runs on | Your machine | A model |
| Deterministic | Yes | No |
| Cost per call | CPU only | Tokens |
| Right for | Arithmetic, database reads, API calls, dates | Summarising, classifying, rewriting, extracting |
To the model, and to the kernel's invocation machinery, they are indistinguishable — both appear in the registry with a name, description and parameter list. That uniformity is why you can compose them freely.
Native functions
1from typing import Annotated2from semantic_kernel.functions import kernel_function3from datetime import datetime, timezone45class PolicyPlugin:6 def __init__(self, repo):7 self._repo = repo89 @kernel_function(10 name="get_policy",11 description=("Look up an insurance policy by its number. Returns the "12 "holder name, product, start date, annual premium in "13 "pounds and current status."),14 )15 def get_policy(16 self,17 policy_number: Annotated[str, "Policy number, 5-8 digits, no spaces"],18 ) -> Annotated[str, "One line per field, or NOT FOUND"]:19 row = self._repo.find(policy_number.strip())20 if row is None:21 return f"NOT FOUND: no policy {policy_number!r}."22 return (f"holder={row.holder}\nproduct={row.product}\n"23 f"start={row.start:%Y-%m-%d}\npremium_gbp={row.premium:.2f}\n"24 f"status={row.status}")2526 @kernel_function(27 name="count_claims",28 description=("Count claims filed against a policy within the last N "29 "months. Use for claims history questions."),30 )31 def count_claims(32 self,33 policy_number: Annotated[str, "Policy number, 5-8 digits"],34 months: Annotated[int, "Look-back window in months, 1-120"] = 24,35 ) -> Annotated[str, "count and total value, or NOT FOUND"]:36 months = max(1, min(120, months))37 claims = self._repo.claims(policy_number.strip(), months)38 if not claims:39 return f"0 claims in the last {months} months."40 total = sum(c.amount for c in claims)41 return (f"{len(claims)} claims in the last {months} months, "42 f"total {total:.2f} GBP.")The Annotated types are not documentation for humans. Semantic Kernel reads them and builds the JSON schema the model sees, so "Policy number, 5-8 digits, no spaces" is literally the text the model consults before deciding what to pass. Leave the annotation off and the model gets a bare string with no guidance, and starts sending "policy 88213" — with the word "policy" in it — to a repository that expects digits.
Note also the clamping: months = max(1, min(120, months)). The schema says 1–120; the model mostly respects it; the code enforces it. Schemas constrain what the model is told, not what it can send.
Registering a plugin
kernel.add_plugin(PolicyPlugin(repo), plugin_name="policy")kernel.add_plugin(MathsPlugin(), plugin_name="maths")Functions are now addressable as policy.get_policy and policy.count_claims. The namespace matters once you have thirty functions: policy.search and documents.search can coexist, and you can grant, revoke or audit access at the plugin level rather than one function at a time.
Prompt functions
1from semantic_kernel.prompt_template import PromptTemplateConfig2from semantic_kernel.functions import KernelArguments34kernel.add_function(5 plugin_name="writer",6 function_name="explain_policy",7 prompt_template_config=PromptTemplateConfig(8 name="explain_policy",9 description=("Rewrite policy details in plain English for a customer "10 "who is not an insurance specialist."),11 template=(12 "Explain these policy details to a customer in at most 80 words. "13 "Use no jargon. Do not add any figure that is not present below.\n\n"14 "{{$details}}"15 ),16 input_variables=[17 {"name": "details", "description": "Raw policy fields",18 "is_required": True},19 ],20 execution_settings={"fast": {"max_tokens": 200}},21 ),22){{$details}} is Semantic Kernel's template syntax for a variable. The execution_settings key names the service — "fast" — so this particular function runs on the cheap model. A summarisation function does not need your most capable model, and routing it explicitly is a decision worth making once rather than paying for on every call.
Invoking
1async def main():2 details = await kernel.invoke(3 plugin_name="policy", function_name="get_policy",4 arguments=KernelArguments(policy_number="88213"),5 )6 plain = await kernel.invoke(7 plugin_name="writer", function_name="explain_policy",8 arguments=KernelArguments(details=str(details)),9 )10 print(plain)1112asyncio.run(main())One native function, one prompt function, called through the same API with the same argument object. Nothing in the calling code reveals which is which. That is the composability the uniform abstraction buys: you can replace a prompt function with a native one — because you found a deterministic way to do it — without touching a caller.
Planners: how the kernel decides what to call
The old way, and why it was replaced
Early Semantic Kernel had explicit planner classes. You handed a goal to a SequentialPlanner, it produced a plan — an XML or JSON document listing functions and arguments — and you executed the plan.
Goal: "Tell me if policy 88213 had claims and what renewal costs"Plan produced: step 1: policy.get_policy(policy_number="88213") -> $details step 2: policy.count_claims(policy_number="88213") -> $claims step 3: maths.multiply(a=$details.premium, b="1.08") -> $renewal step 4: writer.explain_policy(details=$details) -> $answerThe appeal was inspectability: you could show a human the plan before running it. The problems were fatal in practice.
| Problem | Consequence |
|---|---|
| Plan generated before any function runs | Cannot adapt when step 1 returns NOT FOUND |
| Whole plan is one model output | A single malformed field invalidates the entire plan |
| Planning prompt lists every function | Expensive, and quality falls as the registry grows |
| Custom plan format | Fought against models' native tool-calling training |
The adaptivity problem is the decisive one. A plan that says "step 2: count claims for 88213" is worthless if step 1 revealed that 88213 does not exist, but a pre-generated plan has no mechanism to notice.
The current way: automatic function calling
Modern Semantic Kernel drops the separate planning phase. The kernel advertises the registry to the model as native tool definitions, the model calls one function, sees the result, and decides the next call with that result in hand.
1from semantic_kernel.connectors.ai.function_choice_behavior import (2 FunctionChoiceBehavior)3from semantic_kernel.connectors.ai.open_ai import OpenAIChatPromptExecutionSettings4from semantic_kernel.contents import ChatHistory56chat = kernel.get_service("smart")78settings = OpenAIChatPromptExecutionSettings(9 service_id="smart", max_tokens=800,10)11settings.function_choice_behavior = FunctionChoiceBehavior.Auto(12 filters={"included_plugins": ["policy", "maths"]},13 maximum_auto_invoke_attempts=5,14)1516history = ChatHistory()17history.add_system_message(18 "You are an insurance assistant. Never state a premium or claim figure "19 "you did not obtain from a function. If a policy is NOT FOUND, say so "20 "and stop."21)22history.add_user_message(23 "Has policy 88213 had claims in the last two years, and what is the "24 "renewal premium at an 8% increase?"25)2627reply = await chat.get_chat_message_content(28 chat_history=history, settings=settings, kernel=kernel,29)Three arguments in FunctionChoiceBehavior.Auto deserve attention.
| Setting | Effect | Why it matters |
|---|---|---|
Auto() | Model chooses whether and which to call | The normal mode |
Required(filters={"included_functions": ["policy-get_policy"]}) | Forces a specific call | Guarantees a lookup happens before anything else |
NoneInvoke() | Advertises functions but calls none | Dry runs; showing a user what would be called |
filters={"included_plugins": [...]} | Narrows the visible registry | Cost and accuracy — see below |
maximum_auto_invoke_attempts | Caps the call chain | The one setting between you and a runaway loop |
Why the plugin filter is worth real money
Every advertised function costs input tokens on every model call in the chain, not just the first. Suppose each function schema averages 160 tokens.
| Registry advertised | Schema tokens per call | 4-call chain | Cost at an example 3 dollars per million |
|---|---|---|---|
| All 42 functions | 6,720 | 26,880 | 0.081 dollars |
| Two plugins, 9 functions | 1,440 | 5,760 | 0.017 dollars |
A 4.7x difference in schema overhead alone, before a single useful token is generated. Across 200,000 requests a month that is roughly 12,800 dollars. And the narrow registry is also more accurate: with 42 functions the model must discriminate among 42 plausible options, and confusion between policy.search and documents.search becomes routine.
Semantic Kernel next to other framework styles
| Dimension | Semantic Kernel | Chain/agent libraries | Graph runtimes |
|---|---|---|---|
| Central abstraction | Kernel with a function registry | Agent holding a tool list | State machine of nodes |
| Native and model functions | Same abstraction | Separate concepts | Separate concepts |
| Control flow | Model-driven, filtered registry | Model-driven | You define it explicitly |
| Languages | Python, C#, Java | Mostly Python and JS | Mostly Python and JS |
| Cross-cutting hooks | Filters on every invocation | Callbacks | Node wrappers |
| Strongest when | Many existing services; governance and audit matter; mixed-language estate | Rapid prototyping; large ecosystem of integrations | Workflow with branches, cycles and approval gates |
The C# and Java support is not a footnote. If your existing services are .NET, Semantic Kernel lets you expose them as plugins in the language they already live in, with the same registry semantics as the Python side. Its successor, Microsoft Agent Framework, keeps first-class .NET and Python support, which is still rare among agent frameworks.
Every function you advertise costs tokens on every call in the chain, and adds one more thing the model can choose by mistake. Narrow the registry per request.
Where people go wrong
Descriptions written for humans
description="Gets policy" is the commonest defect and the most costly. The description is the model's only basis for choosing. Say what it returns, when to use it, and when not to. A useful test: hand the description alone to a colleague and ask whether they could call the function correctly. If not, neither can the model.
Advertising the whole registry
Covered above with numbers. The default is to expose everything, and the default is wrong at any real scale. Filter by plugin, and route requests to plugin sets with a cheap classification step if you have several domains.
Prompt functions where native ones belong
A prompt function that computes an 8 per cent premium increase costs tokens, adds latency, and gets it wrong occasionally. A native function that multiplies two floats costs nothing and is always right. The rule: if the operation has one correct answer that code can produce, it must be native. Reserve prompt functions for genuine judgement — summarising, classifying, rewriting.
No cap on auto-invoke
maximum_auto_invoke_attempts defaults to 5, and it is tempting to raise it when a long chain gets cut off. Think first: a function that returns an unhelpful message — say, an empty string on a failed lookup — invites the model to retry, and with a high cap it will. Keep the cap small, set it explicitly so the limit is visible in code, and make failure messages say what to do next: "NOT FOUND: no policy 88213. Do not retry; ask the user to confirm the number."
Forgetting that functions are async
kernel.invoke is a coroutine. Calling it without await gives you a coroutine object, which then fails downstream with a confusing type error rather than at the call site. Native functions may be sync or async; the kernel handles both, but the invocation is always awaited.
What this shape is good for
Semantic Kernel earns its place when the functions already exist and the governance questions are real. If your organisation has a hundred internal services, an audit requirement, and a mixed C# and Python estate, the registry model is doing genuine work: it gives you one place to see what the assistant can do, one place to change it, and a filter mechanism that lets you scope capability per request without redeploying.
If you are one developer with four tools and no compliance team, the same machinery is ceremony. A plain tool-calling loop gets you there in a tenth of the code.
When you do build on it, three habits pay for themselves quickly. Write function descriptions as if they were API documentation for an external consumer, because functionally that is what they are. Keep the advertised registry small per request and let a cheap routing step decide which plugins are visible. And put your invariants in the function bodies rather than in the schema — clamp ranges, validate formats, return explicit NOT FOUND messages that tell the model what to do next. The schema shapes the model's intent; only your code decides what actually happens.