AI agents in production: trust, tools, and timeouts
Where agents help, where they fail, and how to contain blast radius when a model calls your APIs.

Shipping an agent demo is easy. Shipping an agent that can touch customer data is a different sport — same models, different blast radius.
The demo path is: prompt, tool call, applause. The production path is: prompt, tool call, audit log, approval, budget trip, pager, and a support ticket that starts with “the AI did what?”
Trust is a product feature
Decide what the agent is allowed to read, propose, and execute. Most teams skip “propose” and jump to execute — then spend the next quarter adding kill switches.
Read widely. Propose clearly. Execute narrowly. If your marketing says “autonomous” but your risk model needs a human, change the marketing — not the safety model.
Tools need contracts
Every tool call should have:
- typed inputs / outputs (the model will invent fields; your API should not)
- idempotency keys where side effects exist
- a human-readable audit trail
If you cannot replay what the agent did last Tuesday, you do not have an operations story. You have a demo with production credentials.
Timeouts beat clever retries
Agents loop. Loops amplify cost and damage. Prefer hard budgets — max tool calls, max wall clock, max tokens — over “one more attempt.”
When the budget trips, stop and hand control back. “One more try” is how incidents get long and invoices get interesting.
Takeaway
Treat agents like untrusted coworkers with API keys: narrow tools, visible actions, and a stop button that actually stops.
The goal is not magic. It is leverage with brakes — and a log someone can read without opening six dashboards.
Found this insightful? Like or share with your team:
Spread good engineering craft & architecture lessons.
Comments
Email is not published. Keep it professional.