The failure pattern
An agent with fifteen tools and a twenty-step ceiling looks powerful in a design document. In production it loops, calls the wrong tool confidently, and burns tokens discovering what it cannot do.
Agents degrade as the decision space grows. Every additional tool is another chance to choose incorrectly at every step.
Narrow the surface
Give the agent the fewest tools that accomplish the task. If two tools are nearly always called together, make them one. If a tool is used in 2% of conversations, consider whether a different entry point should handle that case entirely.
Name and describe tools for the model, not for your codebase. search_knowledge_base with a description explaining exactly when to call it outperforms query_index with a terse one.
Cap the steps
Most production agents do their best work in two to five steps. Set a hard ceiling and define what happens when it is reached — usually escalation to a human, or an honest statement that the agent could not complete the task.
An agent that stops and says so is far better than one that loops until the budget runs out.
Gate consequential actions
Reading is safe. Writing, sending, paying and deleting are not. Anything irreversible or outward-facing belongs behind explicit human approval, regardless of how confident the agent seems.
The engineering question is not "can the agent do this?" but "what happens when it does this wrongly at three in the morning?"
Log everything
Every tool call, input and decision. When an agent behaves strangely — and it will — you need the trace. Teams that skip this spend days reproducing behaviour they could have read directly.
