Course Content
AI Agent Frameworks
4 sections · 15 lessons
Integrating Native Functions and External APIs
A team exposed their whole analytics platform to a model through one function. It looked efficient:
1@kernel_function(name="data_operation", description="Perform a data operation")2def data_operation(self, action: str, params: str) -> str:3 """action: query|aggregate|export|compare ; params: JSON string"""4 ...One function, four capabilities, tidy code. In testing across 50 questions it produced a correct call 29 times. The other 21 failures broke down like this:
| Failure | Count | What the model sent |
|---|---|---|
Invalid action value | 6 | "summarise", "get", "fetch" |
Malformed JSON in params | 8 | Single quotes, trailing commas, unquoted keys |
| Right action, wrong param names | 5 | {"start": "..."} where the code wanted from_date |
| Params for a different action | 2 | {"group_by": ...} passed with action="export" |
58 per cent success. They then split it into four functions with typed parameters and no JSON string anywhere. Same model, same questions: 47 out of 50.
Nothing about the model changed. What changed is what it was asked to produce. Every constraint you push into the schema is one the model sees in a form it reliably follows; every constraint you leave in a docstring is one it can easily miss. Understanding exactly what Semantic Kernel sends the model is therefore not a curiosity — it is the mechanism by which you control accuracy.
What actually reaches the model
Take this function:
1from typing import Annotated2from semantic_kernel.functions import kernel_function34class SalesPlugin:5 @kernel_function(6 name="get_revenue",7 description=("Total revenue for one region over a date range. "8 "Use for questions about how much was sold. "9 "Does not include refunds - use get_refunds for those."),10 )11 def get_revenue(12 self,13 region: Annotated[str, "One of: uk, de, fr, es, nordics"],14 from_date: Annotated[str, "Start date inclusive, ISO format YYYY-MM-DD"],15 to_date: Annotated[str, "End date inclusive, ISO format YYYY-MM-DD"],16 currency: Annotated[str, "Output currency: gbp or eur"] = "gbp",17 ) -> Annotated[str, "Revenue total with currency and row count"]:18 ...Semantic Kernel converts it to this, and this is the entirety of what the model knows about your function:
1{2 "type": "function",3 "function": {4 "name": "sales-get_revenue",5 "description": "Total revenue for one region over a date range. Use for questions about how much was sold. Does not include refunds - use get_refunds for those.",6 "parameters": {7 "type": "object",8 "properties": {9 "region": {"type": "string",10 "description": "One of: uk, de, fr, es, nordics"},11 "from_date": {"type": "string",12 "description": "Start date inclusive, ISO format YYYY-MM-DD"},13 "to_date": {"type": "string",14 "description": "End date inclusive, ISO format YYYY-MM-DD"},15 "currency": {"type": "string",16 "description": "Output currency: gbp or eur"}17 },18 "required": ["region", "from_date", "to_date"]19 }20 }21}Four facts follow directly from that JSON, and each has a practical consequence.
| Fact | Consequence |
|---|---|
| The function body is not sent | Any rule that exists only in code is invisible to the model |
| Type hints become JSON types | int gets integer validation for free; str gets nothing |
| Defaults make a parameter optional | currency is absent from required |
The Annotated string is the only per-parameter guidance | Omit it and the model guesses the format |
If a rule is not in the function description, a parameter description or a type, the model does not know it. The code enforcing it will simply produce errors the model cannot explain.
Push constraints into types
Compare two versions of the same parameter:
1from enum import Enum23class Region(str, Enum):4 UK = "uk"; DE = "de"; FR = "fr"; ES = "es"; NORDICS = "nordics"56# Weak: a free string with advice7region: Annotated[str, "One of: uk, de, fr, es, nordics"]89# Strong: an enum, which becomes {"enum": ["uk","de","fr","es","nordics"]}10region: Annotated[Region, "Sales region"]With the enum, the model sees the exact list of allowed values and almost never strays from it; with strict function calling switched on at the provider, a value such as "United Kingdom" cannot be emitted at all. With the free string it is a live possibility roughly one call in twelve. Either way, keep a check in the function body for the rare call that slips through. The same applies to int versus a string of digits, and to bool versus "true".
Scoping functions well
One job per function
The rewrite that took the analytics team from 29 to 47:
1class SalesPlugin:2 @kernel_function(name="get_revenue", description="Total revenue for one "3 "region over a date range. Does not include refunds.")4 def get_revenue(self, region: Annotated[Region, "Sales region"],5 from_date: Annotated[str, "ISO date YYYY-MM-DD"],6 to_date: Annotated[str, "ISO date YYYY-MM-DD"]) -> str: ...78 @kernel_function(name="compare_regions", description="Compare total revenue "9 "between exactly two regions over the same date range.")10 def compare_regions(self, region_a: Annotated[Region, "First region"],11 region_b: Annotated[Region, "Second region"],12 from_date: Annotated[str, "ISO date YYYY-MM-DD"],13 to_date: Annotated[str, "ISO date YYYY-MM-DD"]) -> str: ...1415 @kernel_function(name="top_products", description="The highest-revenue "16 "products in a region over a date range.")17 def top_products(self, region: Annotated[Region, "Sales region"],18 from_date: Annotated[str, "ISO date YYYY-MM-DD"],19 to_date: Annotated[str, "ISO date YYYY-MM-DD"],20 limit: Annotated[int, "How many, 1-50"] = 10) -> str: ...Longer to write, and worth it. Each function has one meaning, typed parameters, and a description that distinguishes it from its neighbours.
| One god function | Four scoped functions | |
|---|---|---|
| Schema tokens advertised | ~90 | ~520 |
| Correct calls, 50 questions | 29 | 47 |
| Failure mode | Malformed nested JSON | Occasional wrong function choice |
| Recoverable by the model? | Rarely — it repeats the same shape | Usually — a wrong function returns a clear message |
| Auditable | No — every call looks the same in logs | Yes — the log names the operation |
The token column is the honest trade: four scoped functions cost about 430 extra input tokens per model call. On a three-call chain that is 1,290 tokens, or 0.004 dollars at 3 dollars per million. The god function's 21 failures each cost a retry — a whole extra round trip of perhaps 2,500 tokens. At the observed failure rate that is 0.42×2,500=1,050 wasted tokens per question on average, against 1,290 spent deliberately. On tokens alone the two are roughly even; the difference is that the scoped version answers 47 questions correctly instead of 29.
Optional parameters with sensible defaults
Every required parameter is another thing the model must get right. Make anything with an obvious default optional:
1def top_products(2 self,3 region: Annotated[Region, "Sales region"],4 from_date: Annotated[str, "ISO date; defaults to 30 days ago"] = "",5 to_date: Annotated[str, "ISO date; defaults to today"] = "",6 limit: Annotated[int, "How many products, 1-50"] = 10,7) -> str:8 to_d = date.fromisoformat(to_date) if to_date else date.today()9 from_d = date.fromisoformat(from_date) if from_date else to_d - timedelta(days=30)10 limit = max(1, min(50, limit))"Top products in the UK" now works as a single call with one argument. Without the defaults the model must invent two dates, and it will pick something plausible but arbitrary — often the current calendar year, which is rarely what the user meant.
Automatic function calling in practice
1from semantic_kernel.connectors.ai.function_choice_behavior import (2 FunctionChoiceBehavior)3from semantic_kernel.connectors.ai.open_ai import OpenAIChatPromptExecutionSettings4from semantic_kernel.contents import ChatHistory56kernel.add_plugin(SalesPlugin(db), plugin_name="sales")7kernel.add_plugin(MathsPlugin(), plugin_name="maths")89settings = OpenAIChatPromptExecutionSettings(service_id="smart")10settings.function_choice_behavior = FunctionChoiceBehavior.Auto(11 filters={"included_plugins": ["sales", "maths"]},12 maximum_auto_invoke_attempts=6,13)1415history = ChatHistory()16history.add_system_message(17 "You are a sales analytics assistant.\n"18 "- Never state a figure you did not obtain from a function.\n"19 "- Use maths functions for every calculation, including percentages.\n"20 "- Today's date is {today}. Resolve relative dates before calling.\n"21 "- If a function returns a line starting with ERROR, read it and fix the "22 "call once. If it errors again, tell the user what is missing."23)24history.add_user_message("How did UK revenue in Q1 compare with Germany, "25 "and what's the percentage gap?")2627reply = await chat.get_chat_message_content(28 chat_history=history, settings=settings, kernel=kernel)A trace of that request:
call 1 sales-compare_regions(region_a="uk", region_b="de", from_date="2025-01-01", to_date="2025-03-31") -> uk=4,182,500.00 GBP (12,043 rows) de=3,610,200.00 GBP (10,887 rows)call 2 maths-percent_difference(a=4182500, b=3610200) -> 15.85% (a is larger)reply "UK revenue in Q1 was 4,182,500 GBP against Germany's 3,610,200 GBP - the UK was 15.85% higher."Verify the number rather than trusting it: (4,182,500−3,610,200)/3,610,200=572,300/3,610,200=0.15852, so 15.85 per cent. Correct. Had the model computed it in its head, a plausible-looking 15.9 or 13.7 would have been indistinguishable to the reader.
Note that compare_regions handled in one call what would otherwise have been two get_revenue calls. Functions that match the shape of common questions reduce chain length, and each removed link is a whole round trip saved.
Forcing a call
settings.function_choice_behavior = FunctionChoiceBehavior.Required( filters={"included_functions": ["sales-get_revenue"]})Useful when a lookup must happen — a compliance rule that says the assistant may never answer a revenue question from memory. Required guarantees the call; Auto merely makes it likely.
Calling native functions from inside prompts
Semantic Kernel's template language can invoke a function and drop its output into the prompt. This is different from tool calling: it happens before the model sees anything, deterministically.
1from semantic_kernel.core_plugins import TimePlugin23kernel.add_plugin(TimePlugin(), plugin_name="time") # provides {{time.today}}4kernel.add_function(5 plugin_name="reports",6 function_name="weekly_summary",7 prompt_template_config=PromptTemplateConfig(8 name="weekly_summary",9 description="Write a weekly sales summary for one region.",10 template=(11 "Today is {{time.today}}.\n"12 "Revenue data:\n{{sales.get_revenue region=$region "13 "from_date=$start to_date=$end}}\n"14 "Top products:\n{{sales.top_products region=$region "15 "from_date=$start to_date=$end limit=5}}\n\n"16 "Write a 150-word summary for a regional manager. Use only the "17 "figures above. State the single most actionable observation last."18 ),19 input_variables=[20 {"name": "region", "description": "Region code", "is_required": True},21 {"name": "start", "description": "ISO start date", "is_required": True},22 {"name": "end", "description": "ISO end date", "is_required": True},23 ],24 ),25)| Template invocation | Automatic function calling | |
|---|---|---|
| Who decides the call | You, at authoring time | The model, at run time |
| Model calls needed | 1 | 1 per step plus a final one |
| Adapts to the input | No | Yes |
| Cost of a fixed 2-function report | ~1 call | ~3 calls |
| Right for | Known, repeated report shapes | Open-ended questions |
For a weekly report you always produce the same way, the template is three times cheaper and cannot pick the wrong function. Reach for automatic calling only where the sequence genuinely varies.
Defensive error handling
A native function that raises ends the chain. A native function that returns a well-written string keeps the model working. The distinction is worth codifying.
1import requests, time, logging23log = logging.getLogger("sales")45class ExternalAPIPlugin:6 def __init__(self, base, key, timeout=8):7 self.base, self.key, self.timeout = base, key, timeout89 @kernel_function(10 name="get_exchange_rate",11 description=("Current exchange rate between two ISO-4217 currency "12 "codes. Returns a decimal rate, or a line starting with "13 "ERROR: explaining what to do."),14 )15 def get_exchange_rate(16 self,17 base: Annotated[str, "Three-letter ISO code, e.g. GBP"],18 quote: Annotated[str, "Three-letter ISO code, e.g. EUR"],19 ) -> Annotated[str, "Rate or ERROR line"]:20 base, quote = base.upper().strip(), quote.upper().strip()21 if len(base) != 3 or len(quote) != 3:22 return (f"ERROR: {base!r} and {quote!r} must be three-letter ISO "23 f"codes such as GBP, EUR, USD. Fix and call again.")2425 for attempt in range(3):26 try:27 r = requests.get(f"{self.base}/rate",28 params={"base": base, "quote": quote},29 headers={"Authorization": f"Bearer {self.key}"},30 timeout=self.timeout)31 if r.status_code == 429:32 time.sleep(2 ** attempt) # 1s, 2s, 4s33 continue34 if r.status_code in (401, 403):35 log.error("auth failure on rate API")36 raise RuntimeError("exchange rate API credentials invalid")37 if r.status_code == 404:38 return (f"ERROR: no rate for {base}/{quote}. The pair may "39 f"not be supported. Do not retry this pair.")40 r.raise_for_status()41 return f"{base}/{quote} = {r.json()['rate']:.6f}"42 except requests.Timeout:43 log.warning("rate API timeout attempt %d", attempt + 1)44 except requests.ConnectionError:45 log.warning("rate API unreachable attempt %d", attempt + 1)4647 return ("ERROR: exchange rate service unavailable after 3 attempts. "48 "Continue without conversion and say so in your answer.")Four categories, four behaviours:
| Category | Example | Behaviour | Reason |
|---|---|---|---|
| Bad arguments | base="pounds" | Return ERROR: with the correct format | The model can fix it next step |
| Transient | 429, timeout | Retry with exponential backoff inside the function | Cheaper than another model round trip |
| Permanent, request-specific | 404 unsupported pair | Return ERROR: and say "do not retry" | Stops a retry loop the model would otherwise enter |
| Fatal, system-wide | 401 bad key | Raise | No model behaviour fixes a missing credential |
The backoff arithmetic matters. Three attempts at 1, 2 and 4 seconds is 7 seconds of sleeping plus three request timeouts of 8 seconds — a 31-second worst case inside one function. If your assistant has a 30-second response budget, that single function can consume all of it. Either reduce the per-request timeout to 4 seconds or drop to two attempts; the point is to choose deliberately rather than discover it in production.
A constraint expressed as a type is one the model almost never breaks, and one the provider can enforce. The same constraint expressed as a sentence is one it will break roughly one call in twelve.
Six mistakes that cost real accuracy
| Mistake | Symptom | Fix |
|---|---|---|
| Free strings where an enum belongs | Invalid values roughly 1 call in 12 | Enum subclass, which puts the allowed values in the schema |
| JSON passed as a string parameter | Parse errors, quoting problems | Real typed parameters |
| Description says what, not when | Wrong function chosen | Add "use for…" and "do not use for…" |
| Returning raw objects or dicts | Model sees <object at 0x7f…> | Return a formatted string |
| Returning 4,000 rows | Context blown at step 2 | Cap output; return a summary plus a count |
| Silent empty result | Model retries the same call repeatedly | Return "0 rows found for X. Do not retry." |
The row-cap row is worth its own sentence. A function that returns every matching record is a context-window bomb: 4,000 rows at 30 tokens each is 120,000 tokens injected into a conversation that then gets re-sent on every subsequent step. Always cap, always say how many were truncated: "showing 20 of 4,183 rows" lets the model narrow its query instead of guessing.
Designing the function surface
The practical discipline is to design your function surface from the questions users actually ask, not from your database schema. Collect thirty real questions before writing any function. Group them. Each group that a single call could answer becomes a function — which is how compare_regions comes to exist alongside get_revenue, even though it is technically redundant. A function that matches a common question removes an entire round trip.
Then write the descriptions before the bodies. If you cannot describe a function in two sentences that make clear when to use it and when not to, it is doing too much and needs splitting — the description is a design tool, not documentation added afterwards.
Finally, test the schema and not just the code. Keep a fixed set of twenty questions with the expected function name and arguments for each, run it after every change to a description or a type hint, and record the score. Descriptions drift, someone adds a fifth function whose description overlaps a fourth, and accuracy quietly falls from 47 out of 50 to 38 with no error anywhere in your logs. The only way you find out is by measuring it.