AI Agent Frameworks

Semantic Kernel Foundations — Kernel, Plugins, and Planners


An insurance company has 340 internal services. Policy lookup, claims history, premium calculation, document generation, fraud scoring — all of them already exist, all of them audited, all of them with an owner who will not let you reimplement their logic in a prompt.

They want an assistant that answers "Has policy 88213 had any claims in the last two years, and what would the renewal premium be at the new rate?" Answering that means calling three existing services in the right order and doing one calculation.

The naive approach is to describe those services to a model and hope. The problem is not that it will not work — it will, roughly. The problem is that the compliance team has questions: which functions is the assistant allowed to call, who approved that list, what arguments were passed on 14 March, and can we turn off the premium calculator for a week without redeploying?

Semantic Kernel is Microsoft's answer, and its shape follows directly from those questions. Where other frameworks treat tools as a list you hand to an agent, Semantic Kernel treats them as a registry owned by a kernel object: functions are registered into named plugins, the registry is inspectable, and the same registry serves both automatic model-driven calling and direct programmatic invocation.

Most agent frameworks give the model a list of tools. Semantic Kernel gives your application a registry of functions, and lets the model call into it. That inversion is the whole design.

How 340 services reach the modelKernel — the registry and the invokerPlugins — one per existing serviceFunctions — native code, or a promptAutomatic function calling picks and runs
The audited service stays where it is; the kernel only advertises a described entry point to it.

The kernel object

The kernel is a container holding three kinds of thing.

ContentsWhat it isExample
AI servicesChat, embedding or image models, each with a service_idA chat model, a cheaper chat model, an embedder
PluginsNamed collections of callable functionspolicy, claims, maths
Filters and settingsCross-cutting hooks and execution configLogging, auth checks, retry policy
Bash
pip install semantic-kernel
Python
import asyncio, osfrom semantic_kernel import Kernelfrom semantic_kernel.connectors.ai.open_ai import OpenAIChatCompletionkernel = Kernel()kernel.add_service(OpenAIChatCompletion(    service_id="fast",    ai_model_id=os.getenv("FAST_MODEL", "gpt-6-luna"),    api_key=API_KEY,))kernel.add_service(OpenAIChatCompletion(    service_id="smart",    ai_model_id=os.getenv("SMART_MODEL", "gpt-6-sol"),    api_key=API_KEY,))

Two services registered, addressable by id. That is not a convenience feature — it is how you route classification to a model costing a fraction of the price and reserve the expensive one for reasoning. Registering both at kernel level means the choice is made per function, in configuration, rather than baked into code.

Functions: the unit everything is built from

Semantic Kernel has exactly one abstraction for "a thing that can be called", and both flavours share it.

Native functionPrompt function
BodyPython codeA prompt template
Runs onYour machineA model
DeterministicYesNo
Cost per callCPU onlyTokens
Right forArithmetic, database reads, API calls, datesSummarising, classifying, rewriting, extracting

To the model, and to the kernel's invocation machinery, they are indistinguishable — both appear in the registry with a name, description and parameter list. That uniformity is why you can compose them freely.

Native functions

Python
from typing import Annotatedfrom semantic_kernel.functions import kernel_functionfrom datetime import datetime, timezoneclass PolicyPlugin:    def __init__(self, repo):        self._repo = repo    @kernel_function(        name="get_policy",        description=("Look up an insurance policy by its number. Returns the "                     "holder name, product, start date, annual premium in "                     "pounds and current status."),    )    def get_policy(        self,        policy_number: Annotated[str, "Policy number, 5-8 digits, no spaces"],    ) -> Annotated[str, "One line per field, or NOT FOUND"]:        row = self._repo.find(policy_number.strip())        if row is None:            return f"NOT FOUND: no policy {policy_number!r}."        return (f"holder={row.holder}\nproduct={row.product}\n"                f"start={row.start:%Y-%m-%d}\npremium_gbp={row.premium:.2f}\n"                f"status={row.status}")    @kernel_function(        name="count_claims",        description=("Count claims filed against a policy within the last N "                     "months. Use for claims history questions."),    )    def count_claims(        self,        policy_number: Annotated[str, "Policy number, 5-8 digits"],        months: Annotated[int, "Look-back window in months, 1-120"] = 24,    ) -> Annotated[str, "count and total value, or NOT FOUND"]:        months = max(1, min(120, months))        claims = self._repo.claims(policy_number.strip(), months)        if not claims:            return f"0 claims in the last {months} months."        total = sum(c.amount for c in claims)        return (f"{len(claims)} claims in the last {months} months, "                f"total {total:.2f} GBP.")

The Annotated types are not documentation for humans. Semantic Kernel reads them and builds the JSON schema the model sees, so "Policy number, 5-8 digits, no spaces" is literally the text the model consults before deciding what to pass. Leave the annotation off and the model gets a bare string with no guidance, and starts sending "policy 88213" — with the word "policy" in it — to a repository that expects digits.

Note also the clamping: months = max(1, min(120, months)). The schema says 1–120; the model mostly respects it; the code enforces it. Schemas constrain what the model is told, not what it can send.

Registering a plugin

Python
kernel.add_plugin(PolicyPlugin(repo), plugin_name="policy")kernel.add_plugin(MathsPlugin(), plugin_name="maths")

Functions are now addressable as policy.get_policy and policy.count_claims. The namespace matters once you have thirty functions: policy.search and documents.search can coexist, and you can grant, revoke or audit access at the plugin level rather than one function at a time.

Prompt functions

Python
from semantic_kernel.prompt_template import PromptTemplateConfigfrom semantic_kernel.functions import KernelArgumentskernel.add_function(    plugin_name="writer",    function_name="explain_policy",    prompt_template_config=PromptTemplateConfig(        name="explain_policy",        description=("Rewrite policy details in plain English for a customer "                     "who is not an insurance specialist."),        template=(            "Explain these policy details to a customer in at most 80 words. "            "Use no jargon. Do not add any figure that is not present below.\n\n"            "{{$details}}"        ),        input_variables=[            {"name": "details", "description": "Raw policy fields",             "is_required": True},        ],        execution_settings={"fast": {"max_tokens": 200}},    ),)

{{$details}} is Semantic Kernel's template syntax for a variable. The execution_settings key names the service — "fast" — so this particular function runs on the cheap model. A summarisation function does not need your most capable model, and routing it explicitly is a decision worth making once rather than paying for on every call.

Invoking

Python
async def main():    details = await kernel.invoke(        plugin_name="policy", function_name="get_policy",        arguments=KernelArguments(policy_number="88213"),    )    plain = await kernel.invoke(        plugin_name="writer", function_name="explain_policy",        arguments=KernelArguments(details=str(details)),    )    print(plain)asyncio.run(main())

One native function, one prompt function, called through the same API with the same argument object. Nothing in the calling code reveals which is which. That is the composability the uniform abstraction buys: you can replace a prompt function with a native one — because you found a deterministic way to do it — without touching a caller.

Planners: how the kernel decides what to call

The old way, and why it was replaced

Early Semantic Kernel had explicit planner classes. You handed a goal to a SequentialPlanner, it produced a plan — an XML or JSON document listing functions and arguments — and you executed the plan.

Text
Goal: "Tell me if policy 88213 had claims and what renewal costs"Plan produced:  step 1: policy.get_policy(policy_number="88213")     -> $details  step 2: policy.count_claims(policy_number="88213")   -> $claims  step 3: maths.multiply(a=$details.premium, b="1.08") -> $renewal  step 4: writer.explain_policy(details=$details)      -> $answer

The appeal was inspectability: you could show a human the plan before running it. The problems were fatal in practice.

ProblemConsequence
Plan generated before any function runsCannot adapt when step 1 returns NOT FOUND
Whole plan is one model outputA single malformed field invalidates the entire plan
Planning prompt lists every functionExpensive, and quality falls as the registry grows
Custom plan formatFought against models' native tool-calling training

The adaptivity problem is the decisive one. A plan that says "step 2: count claims for 88213" is worthless if step 1 revealed that 88213 does not exist, but a pre-generated plan has no mechanism to notice.

The current way: automatic function calling

Modern Semantic Kernel drops the separate planning phase. The kernel advertises the registry to the model as native tool definitions, the model calls one function, sees the result, and decides the next call with that result in hand.

Python
from semantic_kernel.connectors.ai.function_choice_behavior import (    FunctionChoiceBehavior)from semantic_kernel.connectors.ai.open_ai import OpenAIChatPromptExecutionSettingsfrom semantic_kernel.contents import ChatHistorychat = kernel.get_service("smart")settings = OpenAIChatPromptExecutionSettings(    service_id="smart", max_tokens=800,)settings.function_choice_behavior = FunctionChoiceBehavior.Auto(    filters={"included_plugins": ["policy", "maths"]},    maximum_auto_invoke_attempts=5,)history = ChatHistory()history.add_system_message(    "You are an insurance assistant. Never state a premium or claim figure "    "you did not obtain from a function. If a policy is NOT FOUND, say so "    "and stop.")history.add_user_message(    "Has policy 88213 had claims in the last two years, and what is the "    "renewal premium at an 8% increase?")reply = await chat.get_chat_message_content(    chat_history=history, settings=settings, kernel=kernel,)

Three arguments in FunctionChoiceBehavior.Auto deserve attention.

SettingEffectWhy it matters
Auto()Model chooses whether and which to callThe normal mode
Required(filters={"included_functions": ["policy-get_policy"]})Forces a specific callGuarantees a lookup happens before anything else
NoneInvoke()Advertises functions but calls noneDry runs; showing a user what would be called
filters={"included_plugins": [...]}Narrows the visible registryCost and accuracy — see below
maximum_auto_invoke_attemptsCaps the call chainThe one setting between you and a runaway loop

Why the plugin filter is worth real money

Every advertised function costs input tokens on every model call in the chain, not just the first. Suppose each function schema averages 160 tokens.

Registry advertisedSchema tokens per call4-call chainCost at an example 3 dollars per million
All 42 functions6,72026,8800.081 dollars
Two plugins, 9 functions1,4405,7600.017 dollars

A 4.7x difference in schema overhead alone, before a single useful token is generated. Across 200,000 requests a month that is roughly 12,800 dollars. And the narrow registry is also more accurate: with 42 functions the model must discriminate among 42 plausible options, and confusion between policy.search and documents.search becomes routine.

Semantic Kernel next to other framework styles

DimensionSemantic KernelChain/agent librariesGraph runtimes
Central abstractionKernel with a function registryAgent holding a tool listState machine of nodes
Native and model functionsSame abstractionSeparate conceptsSeparate concepts
Control flowModel-driven, filtered registryModel-drivenYou define it explicitly
LanguagesPython, C#, JavaMostly Python and JSMostly Python and JS
Cross-cutting hooksFilters on every invocationCallbacksNode wrappers
Strongest whenMany existing services; governance and audit matter; mixed-language estateRapid prototyping; large ecosystem of integrationsWorkflow with branches, cycles and approval gates

The C# and Java support is not a footnote. If your existing services are .NET, Semantic Kernel lets you expose them as plugins in the language they already live in, with the same registry semantics as the Python side. Its successor, Microsoft Agent Framework, keeps first-class .NET and Python support, which is still rare among agent frameworks.

Every function you advertise costs tokens on every call in the chain, and adds one more thing the model can choose by mistake. Narrow the registry per request.

Where people go wrong

Descriptions written for humans

description="Gets policy" is the commonest defect and the most costly. The description is the model's only basis for choosing. Say what it returns, when to use it, and when not to. A useful test: hand the description alone to a colleague and ask whether they could call the function correctly. If not, neither can the model.

Advertising the whole registry

Covered above with numbers. The default is to expose everything, and the default is wrong at any real scale. Filter by plugin, and route requests to plugin sets with a cheap classification step if you have several domains.

Prompt functions where native ones belong

A prompt function that computes an 8 per cent premium increase costs tokens, adds latency, and gets it wrong occasionally. A native function that multiplies two floats costs nothing and is always right. The rule: if the operation has one correct answer that code can produce, it must be native. Reserve prompt functions for genuine judgement — summarising, classifying, rewriting.

No cap on auto-invoke

maximum_auto_invoke_attempts defaults to 5, and it is tempting to raise it when a long chain gets cut off. Think first: a function that returns an unhelpful message — say, an empty string on a failed lookup — invites the model to retry, and with a high cap it will. Keep the cap small, set it explicitly so the limit is visible in code, and make failure messages say what to do next: "NOT FOUND: no policy 88213. Do not retry; ask the user to confirm the number."

Forgetting that functions are async

kernel.invoke is a coroutine. Calling it without await gives you a coroutine object, which then fails downstream with a confusing type error rather than at the call site. Native functions may be sync or async; the kernel handles both, but the invocation is always awaited.

What this shape is good for

Semantic Kernel earns its place when the functions already exist and the governance questions are real. If your organisation has a hundred internal services, an audit requirement, and a mixed C# and Python estate, the registry model is doing genuine work: it gives you one place to see what the assistant can do, one place to change it, and a filter mechanism that lets you scope capability per request without redeploying.

If you are one developer with four tools and no compliance team, the same machinery is ceremony. A plain tool-calling loop gets you there in a tenth of the code.

When you do build on it, three habits pay for themselves quickly. Write function descriptions as if they were API documentation for an external consumer, because functionally that is what they are. Keep the advertised registry small per request and let a cheap routing step decide which plugins are visible. And put your invariants in the function bodies rather than in the schema — clamp ranges, validate formats, return explicit NOT FOUND messages that tell the model what to do next. The schema shapes the model's intent; only your code decides what actually happens.