AI Agent Frameworks

Integrating Native Functions and External APIs


A team exposed their whole analytics platform to a model through one function. It looked efficient:

Python
@kernel_function(name="data_operation", description="Perform a data operation")def data_operation(self, action: str, params: str) -> str:    """action: query|aggregate|export|compare ; params: JSON string"""    ...

One function, four capabilities, tidy code. In testing across 50 questions it produced a correct call 29 times. The other 21 failures broke down like this:

FailureCountWhat the model sent
Invalid action value6"summarise", "get", "fetch"
Malformed JSON in params8Single quotes, trailing commas, unquoted keys
Right action, wrong param names5{"start": "..."} where the code wanted from_date
Params for a different action2{"group_by": ...} passed with action="export"

58 per cent success. They then split it into four functions with typed parameters and no JSON string anywhere. Same model, same questions: 47 out of 50.

Nothing about the model changed. What changed is what it was asked to produce. Every constraint you push into the schema is one the model sees in a form it reliably follows; every constraint you leave in a docstring is one it can easily miss. Understanding exactly what Semantic Kernel sends the model is therefore not a curiosity — it is the mechanism by which you control accuracy.

One catch-all function, or a scoped surfaceThe whole platform in one function• Arguments are free-form strings• Description cannot say what is valid• Every misuse looks like a valid call• Errors surface far from the causeOne job per function• Types push constraints to the boundary• Optional parameters carry defaults• Description names one clear job• Wrong calls fail before they execute
What reaches the model is the signature and the description — a type is a constraint you do not have to explain in prose.

What actually reaches the model

Take this function:

Python
from typing import Annotatedfrom semantic_kernel.functions import kernel_functionclass SalesPlugin:    @kernel_function(        name="get_revenue",        description=("Total revenue for one region over a date range. "                     "Use for questions about how much was sold. "                     "Does not include refunds - use get_refunds for those."),    )    def get_revenue(        self,        region: Annotated[str, "One of: uk, de, fr, es, nordics"],        from_date: Annotated[str, "Start date inclusive, ISO format YYYY-MM-DD"],        to_date: Annotated[str, "End date inclusive, ISO format YYYY-MM-DD"],        currency: Annotated[str, "Output currency: gbp or eur"] = "gbp",    ) -> Annotated[str, "Revenue total with currency and row count"]:        ...

Semantic Kernel converts it to this, and this is the entirety of what the model knows about your function:

JSON
{  "type": "function",  "function": {    "name": "sales-get_revenue",    "description": "Total revenue for one region over a date range. Use for questions about how much was sold. Does not include refunds - use get_refunds for those.",    "parameters": {      "type": "object",      "properties": {        "region":    {"type": "string",                      "description": "One of: uk, de, fr, es, nordics"},        "from_date": {"type": "string",                      "description": "Start date inclusive, ISO format YYYY-MM-DD"},        "to_date":   {"type": "string",                      "description": "End date inclusive, ISO format YYYY-MM-DD"},        "currency":  {"type": "string",                      "description": "Output currency: gbp or eur"}      },      "required": ["region", "from_date", "to_date"]    }  }}

Four facts follow directly from that JSON, and each has a practical consequence.

FactConsequence
The function body is not sentAny rule that exists only in code is invisible to the model
Type hints become JSON typesint gets integer validation for free; str gets nothing
Defaults make a parameter optionalcurrency is absent from required
The Annotated string is the only per-parameter guidanceOmit it and the model guesses the format

If a rule is not in the function description, a parameter description or a type, the model does not know it. The code enforcing it will simply produce errors the model cannot explain.

Push constraints into types

Compare two versions of the same parameter:

Python
from enum import Enumclass Region(str, Enum):    UK = "uk"; DE = "de"; FR = "fr"; ES = "es"; NORDICS = "nordics"# Weak: a free string with adviceregion: Annotated[str, "One of: uk, de, fr, es, nordics"]# Strong: an enum, which becomes {"enum": ["uk","de","fr","es","nordics"]}region: Annotated[Region, "Sales region"]

With the enum, the model sees the exact list of allowed values and almost never strays from it; with strict function calling switched on at the provider, a value such as "United Kingdom" cannot be emitted at all. With the free string it is a live possibility roughly one call in twelve. Either way, keep a check in the function body for the rare call that slips through. The same applies to int versus a string of digits, and to bool versus "true".

Scoping functions well

One job per function

The rewrite that took the analytics team from 29 to 47:

Python
class SalesPlugin:    @kernel_function(name="get_revenue", description="Total revenue for one "        "region over a date range. Does not include refunds.")    def get_revenue(self, region: Annotated[Region, "Sales region"],                    from_date: Annotated[str, "ISO date YYYY-MM-DD"],                    to_date: Annotated[str, "ISO date YYYY-MM-DD"]) -> str: ...    @kernel_function(name="compare_regions", description="Compare total revenue "        "between exactly two regions over the same date range.")    def compare_regions(self, region_a: Annotated[Region, "First region"],                        region_b: Annotated[Region, "Second region"],                        from_date: Annotated[str, "ISO date YYYY-MM-DD"],                        to_date: Annotated[str, "ISO date YYYY-MM-DD"]) -> str: ...    @kernel_function(name="top_products", description="The highest-revenue "        "products in a region over a date range.")    def top_products(self, region: Annotated[Region, "Sales region"],                     from_date: Annotated[str, "ISO date YYYY-MM-DD"],                     to_date: Annotated[str, "ISO date YYYY-MM-DD"],                     limit: Annotated[int, "How many, 1-50"] = 10) -> str: ...

Longer to write, and worth it. Each function has one meaning, typed parameters, and a description that distinguishes it from its neighbours.

One god functionFour scoped functions
Schema tokens advertised~90~520
Correct calls, 50 questions2947
Failure modeMalformed nested JSONOccasional wrong function choice
Recoverable by the model?Rarely — it repeats the same shapeUsually — a wrong function returns a clear message
AuditableNo — every call looks the same in logsYes — the log names the operation

The token column is the honest trade: four scoped functions cost about 430 extra input tokens per model call. On a three-call chain that is 1,290 tokens, or 0.004 dollars at 3 dollars per million. The god function's 21 failures each cost a retry — a whole extra round trip of perhaps 2,500 tokens. At the observed failure rate that is 0.42×2,500=1,0500.42 \times 2{,}500 = 1{,}050 wasted tokens per question on average, against 1,290 spent deliberately. On tokens alone the two are roughly even; the difference is that the scoped version answers 47 questions correctly instead of 29.

Optional parameters with sensible defaults

Every required parameter is another thing the model must get right. Make anything with an obvious default optional:

Python
def top_products(    self,    region: Annotated[Region, "Sales region"],    from_date: Annotated[str, "ISO date; defaults to 30 days ago"] = "",    to_date: Annotated[str, "ISO date; defaults to today"] = "",    limit: Annotated[int, "How many products, 1-50"] = 10,) -> str:    to_d = date.fromisoformat(to_date) if to_date else date.today()    from_d = date.fromisoformat(from_date) if from_date else to_d - timedelta(days=30)    limit = max(1, min(50, limit))

"Top products in the UK" now works as a single call with one argument. Without the defaults the model must invent two dates, and it will pick something plausible but arbitrary — often the current calendar year, which is rarely what the user meant.

Automatic function calling in practice

Python
from semantic_kernel.connectors.ai.function_choice_behavior import (    FunctionChoiceBehavior)from semantic_kernel.connectors.ai.open_ai import OpenAIChatPromptExecutionSettingsfrom semantic_kernel.contents import ChatHistorykernel.add_plugin(SalesPlugin(db), plugin_name="sales")kernel.add_plugin(MathsPlugin(), plugin_name="maths")settings = OpenAIChatPromptExecutionSettings(service_id="smart")settings.function_choice_behavior = FunctionChoiceBehavior.Auto(    filters={"included_plugins": ["sales", "maths"]},    maximum_auto_invoke_attempts=6,)history = ChatHistory()history.add_system_message(    "You are a sales analytics assistant.\n"    "- Never state a figure you did not obtain from a function.\n"    "- Use maths functions for every calculation, including percentages.\n"    "- Today's date is {today}. Resolve relative dates before calling.\n"    "- If a function returns a line starting with ERROR, read it and fix the "    "call once. If it errors again, tell the user what is missing.")history.add_user_message("How did UK revenue in Q1 compare with Germany, "                         "and what's the percentage gap?")reply = await chat.get_chat_message_content(    chat_history=history, settings=settings, kernel=kernel)

A trace of that request:

Text
call 1  sales-compare_regions(region_a="uk", region_b="de",                              from_date="2025-01-01", to_date="2025-03-31")        -> uk=4,182,500.00 GBP (12,043 rows)           de=3,610,200.00 GBP (10,887 rows)call 2  maths-percent_difference(a=4182500, b=3610200)        -> 15.85% (a is larger)reply   "UK revenue in Q1 was 4,182,500 GBP against Germany's 3,610,200 GBP -         the UK was 15.85% higher."

Verify the number rather than trusting it: (4,182,500−3,610,200)/3,610,200=572,300/3,610,200=0.15852(4{,}182{,}500 - 3{,}610{,}200) / 3{,}610{,}200 = 572{,}300 / 3{,}610{,}200 = 0.15852, so 15.85 per cent. Correct. Had the model computed it in its head, a plausible-looking 15.9 or 13.7 would have been indistinguishable to the reader.

Note that compare_regions handled in one call what would otherwise have been two get_revenue calls. Functions that match the shape of common questions reduce chain length, and each removed link is a whole round trip saved.

Forcing a call

Python
settings.function_choice_behavior = FunctionChoiceBehavior.Required(    filters={"included_functions": ["sales-get_revenue"]})

Useful when a lookup must happen — a compliance rule that says the assistant may never answer a revenue question from memory. Required guarantees the call; Auto merely makes it likely.

Calling native functions from inside prompts

Semantic Kernel's template language can invoke a function and drop its output into the prompt. This is different from tool calling: it happens before the model sees anything, deterministically.

Python
from semantic_kernel.core_plugins import TimePluginkernel.add_plugin(TimePlugin(), plugin_name="time")   # provides {{time.today}}kernel.add_function(    plugin_name="reports",    function_name="weekly_summary",    prompt_template_config=PromptTemplateConfig(        name="weekly_summary",        description="Write a weekly sales summary for one region.",        template=(            "Today is {{time.today}}.\n"            "Revenue data:\n{{sales.get_revenue region=$region "            "from_date=$start to_date=$end}}\n"            "Top products:\n{{sales.top_products region=$region "            "from_date=$start to_date=$end limit=5}}\n\n"            "Write a 150-word summary for a regional manager. Use only the "            "figures above. State the single most actionable observation last."        ),        input_variables=[            {"name": "region", "description": "Region code", "is_required": True},            {"name": "start", "description": "ISO start date", "is_required": True},            {"name": "end", "description": "ISO end date", "is_required": True},        ],    ),)
Template invocationAutomatic function calling
Who decides the callYou, at authoring timeThe model, at run time
Model calls needed11 per step plus a final one
Adapts to the inputNoYes
Cost of a fixed 2-function report~1 call~3 calls
Right forKnown, repeated report shapesOpen-ended questions

For a weekly report you always produce the same way, the template is three times cheaper and cannot pick the wrong function. Reach for automatic calling only where the sequence genuinely varies.

Defensive error handling

A native function that raises ends the chain. A native function that returns a well-written string keeps the model working. The distinction is worth codifying.

Python
import requests, time, logginglog = logging.getLogger("sales")class ExternalAPIPlugin:    def __init__(self, base, key, timeout=8):        self.base, self.key, self.timeout = base, key, timeout    @kernel_function(        name="get_exchange_rate",        description=("Current exchange rate between two ISO-4217 currency "                     "codes. Returns a decimal rate, or a line starting with "                     "ERROR: explaining what to do."),    )    def get_exchange_rate(        self,        base: Annotated[str, "Three-letter ISO code, e.g. GBP"],        quote: Annotated[str, "Three-letter ISO code, e.g. EUR"],    ) -> Annotated[str, "Rate or ERROR line"]:        base, quote = base.upper().strip(), quote.upper().strip()        if len(base) != 3 or len(quote) != 3:            return (f"ERROR: {base!r} and {quote!r} must be three-letter ISO "                    f"codes such as GBP, EUR, USD. Fix and call again.")        for attempt in range(3):            try:                r = requests.get(f"{self.base}/rate",                                 params={"base": base, "quote": quote},                                 headers={"Authorization": f"Bearer {self.key}"},                                 timeout=self.timeout)                if r.status_code == 429:                    time.sleep(2 ** attempt)          # 1s, 2s, 4s                    continue                if r.status_code in (401, 403):                    log.error("auth failure on rate API")                    raise RuntimeError("exchange rate API credentials invalid")                if r.status_code == 404:                    return (f"ERROR: no rate for {base}/{quote}. The pair may "                            f"not be supported. Do not retry this pair.")                r.raise_for_status()                return f"{base}/{quote} = {r.json()['rate']:.6f}"            except requests.Timeout:                log.warning("rate API timeout attempt %d", attempt + 1)            except requests.ConnectionError:                log.warning("rate API unreachable attempt %d", attempt + 1)        return ("ERROR: exchange rate service unavailable after 3 attempts. "                "Continue without conversion and say so in your answer.")

Four categories, four behaviours:

CategoryExampleBehaviourReason
Bad argumentsbase="pounds"Return ERROR: with the correct formatThe model can fix it next step
Transient429, timeoutRetry with exponential backoff inside the functionCheaper than another model round trip
Permanent, request-specific404 unsupported pairReturn ERROR: and say "do not retry"Stops a retry loop the model would otherwise enter
Fatal, system-wide401 bad keyRaiseNo model behaviour fixes a missing credential

The backoff arithmetic matters. Three attempts at 1, 2 and 4 seconds is 7 seconds of sleeping plus three request timeouts of 8 seconds — a 31-second worst case inside one function. If your assistant has a 30-second response budget, that single function can consume all of it. Either reduce the per-request timeout to 4 seconds or drop to two attempts; the point is to choose deliberately rather than discover it in production.

A constraint expressed as a type is one the model almost never breaks, and one the provider can enforce. The same constraint expressed as a sentence is one it will break roughly one call in twelve.

Six mistakes that cost real accuracy

MistakeSymptomFix
Free strings where an enum belongsInvalid values roughly 1 call in 12Enum subclass, which puts the allowed values in the schema
JSON passed as a string parameterParse errors, quoting problemsReal typed parameters
Description says what, not whenWrong function chosenAdd "use for…" and "do not use for…"
Returning raw objects or dictsModel sees <object at 0x7f…>Return a formatted string
Returning 4,000 rowsContext blown at step 2Cap output; return a summary plus a count
Silent empty resultModel retries the same call repeatedlyReturn "0 rows found for X. Do not retry."

The row-cap row is worth its own sentence. A function that returns every matching record is a context-window bomb: 4,000 rows at 30 tokens each is 120,000 tokens injected into a conversation that then gets re-sent on every subsequent step. Always cap, always say how many were truncated: "showing 20 of 4,183 rows" lets the model narrow its query instead of guessing.

Designing the function surface

The practical discipline is to design your function surface from the questions users actually ask, not from your database schema. Collect thirty real questions before writing any function. Group them. Each group that a single call could answer becomes a function — which is how compare_regions comes to exist alongside get_revenue, even though it is technically redundant. A function that matches a common question removes an entire round trip.

Then write the descriptions before the bodies. If you cannot describe a function in two sentences that make clear when to use it and when not to, it is doing too much and needs splitting — the description is a design tool, not documentation added afterwards.

Finally, test the schema and not just the code. Keep a fixed set of twenty questions with the expected function name and arguments for each, run it after every change to a description or a type hint, and record the score. Descriptions drift, someone adds a fifth function whose description overlaps a fourth, and accuracy quietly falls from 47 out of 50 to 38 with no error anywhere in your logs. The only way you find out is by measuring it.