Course Content
Agents & Tools Interview Prep
6 sections · 40 lessons
What is the difference between tool-calling agents and code-generating agents?
What you need to know
Tool calling
- Fixed menu of typed functions
- One model round trip per step
- Every result enters the context
- Easy to validate, approve and log
Code generation
- Model writes a script
- One script can make dozens of calls
- Only the final output enters the context
- Needs a real sandbox; harder to audit
Why code can be much cheaper
Suppose the agent must find which of 200 customers had more than 3 failed payments. With plain tool calling, that is 200 get_payments calls, each result passing through the model's context. With code, the model writes:
1flagged = []2for cid in list_customer_ids(segment="premium"):3 fails = [p for p in get_payments(cid) if p["status"] == "failed"]4 if len(fails) > 3:5 flagged.append({"customer": cid, "failed": len(fails)})6print(flagged)The loop runs in the sandbox; the model sees only the short flagged list. Research on "code as actions" (for example the CodeAct paper, 2024) and frameworks like smolagents' CodeAgent are built on this idea.
Programmatic tool calling
Some APIs now support a hybrid. In Anthropic's programmatic tool calling, you enable the code execution tool and mark your own tools with allowed_callers so code can call them. The model writes a script; each time the script calls your tool, the API pauses, your code runs the tool, and the result goes back to the script, not into the model's context. You keep your tool definitions and handlers, but gain loops and filtering.
What code generation demands
Generated code is untrusted. Run it in a proper sandbox:
- No network, or an allowlist of domains.
- No credentials in the environment; tools that need secrets stay on your side.
- CPU, memory and time limits.
- A throwaway filesystem.
And accept that auditing a script is harder than auditing one create_refund(order_id, amount) call.
A real-life example
A SQL analytics agent at an e-commerce company must answer "Which 10 sellers had the biggest rise in return rate from July to August?"
As a tool-calling agent, it ran run_query for July, got 4,200 rows back into its context (about 60,000 tokens), did the same for August, then tried to compute differences in its head — slow, expensive and error-prone.
As a code-generating agent, it wrote one script: two read-only queries, a pandas join, a sort, head(10). Only a 10-row table reached the model. Tokens fell by about 95% and the numbers were exactly right.
Their bank-transfer tool, however, stays a plain tool call: create_transfer has typed arguments, a limit check and a confirmation screen. Nobody wants money moved by an unreviewed script.
Follow-up questions to expect
- "Isn't a code agent just one tool called
run_python?" — Technically yes, but the design question is what the code may reach. With programmatic tool calling, code reaches your tools through your handlers, which keep their checks. - "How do you sandbox generated code?" — Containers or micro-VMs (gVisor, Firecracker), no ambient credentials, egress blocked, resource limits, and a fresh environment per task.
- "Which is more reliable for small models?" — Usually tool calling with strict schemas; writing correct code is a harder skill for small models.