Human-in-the-loop agents that do not fake autonomy
Ship agents that propose, wait, and execute with an audit trail — so “AI in production” means controlled leverage, not unsupervised side effects.

Product demos love full autonomy: the agent plans, calls tools, and “just handles it.” Production teams learn a different lesson the first time an agent refunds the wrong invoice or emails the wrong customer.
Autonomy is not a virtue by itself. Leverage with brakes is.
Three privilege levels — keep them separate
Most agent failures come from collapsing these into one mode:
- Read — inspect tickets, docs, metrics, account state
- Propose — draft a change, a reply, a config diff for a human
- Execute — mutate money, access, messages, or infrastructure
Ship Read freely. Ship Propose with a review UX. Ship Execute only behind explicit approval, scoped tools, and an audit log someone can replay.
If your product copy says “autonomous” but your risk model needs a human, change the copy — not the safety model.
Design the approval as a product surface
A yes/no modal bolted on at the end is how approvals get rubber-stamped. Better:
- Show what will change in human language (and the raw tool payload for power users)
- Show blast radius — which accounts, environments, or dollar amounts
- Require re-auth or step-up for irreversible actions
- Record who approved, with timestamp and agent run id
The approval screen is part of the agent. Treat it like a checkout flow, not a CAPTCHA for robots.
Budgets beat cleverness
Agents loop. Loops burn money and amplify mistakes. Give every run hard ceilings:
- max tool calls
- max wall-clock time
- max tokens / cost
- allow-list of tools for that workflow
When the budget trips, stop and hand control back. “One more try” is how incidents get long.
Evaluation is not a vibe check
Before widening execute privileges, keep a regression set of scenarios: happy path, hostile input, partial tool failure, duplicate submission. Run it on every prompt or tool change.
If you cannot say what broke between last week and this week, you do not have an agent product — you have a demo with production credentials.
Takeaway
Customers do not pay for the fantasy of unsupervised AI. They pay for outcomes with contained risk.
Build agents that read widely, propose clearly, and execute narrowly — with a human in the loop until the evidence says otherwise.
Found this insightful? Like or share with your team:
Spread good engineering craft & architecture lessons.
Comments
Email is not published. Keep it professional.