Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
What are best practices when defining tools (inputs, outputs, descriptions)?
What you need to know
Names
- Verb first, specific, snake_case:
search_flights,get_booking,cancel_booking. - One tool, one job. Avoid
manage_booking(action=...)with ten modes — the model must then get both the tool and the mode right. - Group related tools with a prefix (
github_list_issues,github_add_label), especially when tools come from several sources.
Descriptions
Say what, when, when not, and what comes back:
Search available hotels for given dates. Use when the user wants hotel optionsor prices. Do not use to check an existing reservation - use get_booking.Returns at most 10 hotels sorted by price, each with hotel_id, name, area,price_per_night_inr and rating.Current Claude models are fairly conservative about calling tools, and trigger conditions ("use when…") in the description measurably improve how often they call the right one.
Input schemas
| Problem | Weak schema | Better schema |
|---|---|---|
| Unclear format | "date": {"type": "string"} | "format": "date", description "YYYY-MM-DD, e.g. 2026-12-20" |
| Unclear units | "budget": {"type": "number"} | "budget_inr" with "total for the stay, in rupees" |
| Free text for a fixed set | "cabin": {"type": "string"} | "enum": ["economy", "premium_economy", "business"] |
| Invented IDs | "hotel_id" with no hint | "hotel_id from search_hotels results, e.g. HTL-4411" |
| Extra junk fields | nothing | "additionalProperties": false |
Strict mode
Constrained decoding makes the model's arguments always match the schema.
- Anthropic: add
"strict": trueat the top level of the tool definition (next toname), withadditionalProperties: falseand arequiredlist. - OpenAI:
"strict": trueon the function tool. Every property must be listed inrequired; an optional field is written as nullable, e.g."type": ["string", "null"].
Strict mode fixes shape errors (missing fields, wrong types, bad enum values). It does not fix meaning errors (a real but wrong date), so you still validate in the handler.
Outputs
- Return only the fields the model needs, not a 40-field database row.
- Cap and paginate: "10 results; call again with
page=2for more". - Use stable IDs the model can pass to the next tool.
- Prefer structured JSON or clean text over raw HTML.
Errors
Make errors instructive: "No hotel with id HTL-44. Use a hotel_id from search_hotels results." beats "404". The model reads the error and usually fixes the call on the next turn.
How many tools
Accuracy falls as the menu grows and overlaps. If you have dozens of tools, load them by task, merge near-duplicates, or use tool search: mark rarely used tools with defer_loading: true so the model searches for them instead of reading every schema up front (both Anthropic and OpenAI support this in 2026).
A real-life example
A travel startup's first version had one tool:
{"name": "travel_api", "description": "Travel operations", "input_schema": {"type": "object", "properties": { "op": {"type": "string"}, "params": {"type": "object"}}}}The model guessed operation names ("find_flight", "flightSearch"), sent dates as "20/12", and invented hotel IDs. About 1 in 5 calls failed.
They split it into search_flights, search_hotels, get_booking and hold_booking, each with typed parameters, formats, enums and "use when" descriptions, turned on strict mode, and trimmed hotel results from 60 fields to 7. Failed calls dropped to under 1 in 50, and the average tool result shrank from about 9,000 tokens to 900 — which also cut cost per conversation by more than half.
Follow-up questions to expect
- "Should you give examples inside tool definitions?" — Yes, for complex inputs. A format example in the parameter description is cheap; Anthropic also supports example inputs on the tool definition.
- "Should a tool return errors as exceptions or results?" — As results with an error flag, so the model can adapt. Exceptions should only escape for bugs in your harness.
- "How do you know your descriptions are good?" — Build a labelled set of requests with the correct tool and arguments, and measure tool-call accuracy after each change.