An agent is not a chatbot with better prompts. It is an autonomous system that runs a multi-step workflow on your behalf, judges when the work is done, self-corrects or stops, and calls tools inside guardrails you set. AI Central's Ultimate Guide To Building Agents puts a qualifying test before any of that, and says plainly that if your workflow does not clear it, a deterministic solution may suffice. Everything after the test is model, tool and instruction work.
Reviewed 7 August 2026.
The test that comes before the build
Most agent projects fail before a line of code, because the workflow chosen never needed one. AI Central's Ultimate Guide To Building Agents opens there rather than on architecture, arguing agents are uniquely suited to workflows where traditional deterministic and rule-based approaches fall short. Its example is payment fraud analysis.
A traditional rules engine works like a checklist, flagging transactions based on preset criteria. In contrast, an LLM agent functions more like a seasoned investigator, evaluating context, considering subtle patterns, and identifying suspicious activity even when clear-cut rules aren’t violated.
A checklist cannot flag what it was never told to look for. An investigator can. If your problem really is a checklist, you do not have an agent problem.
Three conditions qualify a workflow, each with an example.
- Complex decision-making, meaning nuanced judgment, exceptions or context-sensitive decisions. The example is refund approval.
- Difficult-to-maintain rules, meaning rulesets so extensive and intricate that updates are costly or error-prone. The example is vendor security reviews.
- Heavy reliance on unstructured data, meaning natural language, documents, or conversational interaction. The example is processing a home insurance claim.
And the line most agent content leaves out, from AI Central.
Before committing to building an agent, validate that your use case can meet these criteria clearly. Otherwise, a deterministic solution may suffice
Three components, and only three
Stripped down, an agent is a model powering reasoning and decision-making, tools that are external functions or APIs it uses to take action, and instructions that are the explicit guidelines and guardrails defining behaviour. The code shown uses OpenAI's Agents SDK, and the guide is explicit that the same concepts work in any library or from scratch. The pattern is portable, the SDK is not the point.
On model choice, AI Central refuses the usual reflex.
Not every task requires the smartest model - a simple retrieval or intent classification task may be handled by a smaller, faster model, while harder tasks like deciding whether to approve a refund may benefit from a more capable model
The sequencing is the part to copy. Prototype with the most capable model on every task to establish a baseline, then swap in smaller ones to see whether results hold, so you never prematurely limit the agent and can see where small models fail. Evals first, accuracy target second, cost and latency third. Note the order. Downgrading a model without a baseline is guesswork dressed as cost control.
Tools are an interface problem, not a feature list
Tools come in three types. Data tools retrieve context, querying transaction databases or CRMs, reading PDFs, searching the web. Action tools change something, sending emails, updating a CRM record, handing a ticket to a human. Orchestration tools are other agents, since an agent can be a tool for another.
Two details are easy to skim and expensive to miss. For legacy systems without APIs, agents can rely on computer-use models to drive those applications through web and application UIs, just as a human would, which removes the usual blocker, that the system holding the valuable work has no integration surface. And every tool should carry a standardized definition, enabling many-to-many relationships between tools and agents.
Tools are shared infrastructure. Treat them as one-off functions bolted to one agent and the second agent becomes a rewrite.
Instructions carry most of the failure rate
Clear instructions reduce ambiguity and improve decision-making, which shows up as smoother execution and fewer errors. Four practices carry it.
- Use existing documents. Turn operating procedures, support scripts and policy documents into LLM-friendly routines. In customer service, routines roughly map to individual knowledge base articles.
- Prompt agents to break down tasks. Smaller, clearer steps pulled from dense resources minimise ambiguity.
- Define clear actions. Every step should map to a specific action or output, such as asking for an order number or calling an API for account details, down to the wording of a user-facing message.
- Capture edge cases. Anticipate incomplete information and unexpected questions with conditional steps or branches.
It suggests generating those instructions with advanced models, naming o1 and o3-mini, and supplies the conversion prompt. The structural insight is that your help centre is already a specification, written for humans, and converting it is cheap.
Start with one agent, and keep it there
On orchestration the guide is blunt. The temptation is to immediately build a fully autonomous agent with complex architecture, while customers typically achieve greater success with an incremental approach. Single-agent systems run one model with appropriate tools and instructions in a loop, multi-agent systems spread execution across coordinated agents.
The loop is load-bearing. Every orchestration approach needs the concept of a run, a loop that lets the agent operate until an exit condition is reached, commonly a tool call, a structured output, an error, or a maximum number of turns. In the Agents SDK it continues until a final-output tool fires or the model replies with no tool calls.
Before splitting anything, there is a cheaper lever. Rather than maintaining numerous prompts for distinct use cases, use one flexible base prompt that accepts policy variables, so new cases mean updating variables instead of rewriting workflows. Its own example is a call centre agent greeting a member by name and tenure.
When to split, and the tool-count trap
AI Central's position on adding agents is restrained.
Our general recommendation is to maximize a single agent’s capabilities first.
More agents buy intuitive separation of concepts and cost complexity and overhead. The signal to split is when agents fail to follow complicated instructions or consistently select incorrect tools. Two triggers qualify. Complex logic, where prompts fill with conditional branches and templates stop scaling. And tool overload, where the guide is sharpest.
The issue isn’t solely the number of tools, but their similarity or overlap. Some implementations successfully manage more than 15 well-defined, distinct tools while others struggle with fewer than 10 overlapping tools.
That inverts the usual advice. The count is not the variable, overlap is. The prescribed order is to improve tool clarity first with descriptive names, clear parameters and detailed descriptions, and to split only when that fails. Most teams do the reverse and spin up an agent to escape a naming problem.
Manager or decentralized
Two patterns cover most cases. A central manager agent coordinates specialized agents through tool calls and synthesises the results, which suits workflows where exactly one agent should control execution and hold the user conversation. In the decentralized pattern, agents operate as peers and hand off by specialisation.
Handoff mechanics matter. A handoff is a one-way transfer, and in the Agents SDK it is a type of tool, so calling it immediately starts execution on the new agent while transferring the latest conversation state. The example is a triage agent routing to technical support, sales or order management.
Both patterns are graphs with agents as nodes. In the manager pattern the edges are tool calls, in the decentralized pattern they are handoffs that transfer execution. That distinction decides who is still holding the user when something breaks.
AI Central closes on the same note regardless of pattern.
Regardless of the orchestration pattern, the same principles apply: keep components flexible, composable, and driven by clear, well-structured prompts
What to actually do this week
If you have never shipped an agent, pick a workflow that fails the rules-engine test, convert an existing policy document into a numbered routine, and ship one agent with three tools and a run loop.
If one is already in production and misbehaving, audit tool overlap before adding a second. Rename, tighten parameters, rewrite descriptions, measure. Clarity first, architecture second.
If you are scaling past one agent, choose by ownership, not elegance. One agent keeping the user relationship means the manager pattern. Specialists taking over outright means handoffs.
What is the difference between an AI agent and a chatbot?
Autonomy over multi-step work. Agents execute multi-step workflows on a user's behalf, unlike basic LLM apps or chatbots that only assist or respond. They manage decisions and progress, know when tasks are complete, can self-correct or stop, and use tools within defined guardrails.
When should I not build an agent?
When your workflow does not clearly meet the three criteria of complex decision-making, difficult-to-maintain rules, and heavy reliance on unstructured data. Validate against them first, because otherwise a deterministic solution may suffice.
How many tools can one agent handle?
There is no fixed ceiling. Some implementations manage more than 15 well-defined, distinct tools, while others struggle with fewer than 10 overlapping ones. Overlap decides it, not the raw count, so better names, parameters and descriptions are the first fix.
Which model should I use?
Prototype with the most capable model to set a baseline, then swap in smaller ones where results hold. Simple retrieval or intent classification suits a smaller, faster model, approving a refund does not.