A few months back, a friend’s AI agent sent the same refund confirmation to a customer four times in one afternoon. The model was fine. The prompt was fine. What was missing was any memory of what it had already done and a rule about when to stop. That gap between a slick demo and something you’d trust with a live inbox is where most builds quietly die.
So here’s the build order I use now, and the one I’d hand to anyone trying to make AI agents that survive contact with real users. None of it is exotic. It’s just specific, and specificity is the part most tutorials skip.
Start With the Job, Not the Model
Before you open a single tool, finish this sentence on paper: When ___ happens, the agent should ___, and I’ll know it worked when ___.
Here’s a real one from a small e-commerce team. When a customer emails asking about a refund, the agent should look up the order, check it against the refund policy, and leave a draft reply in the shared inbox. I’ll know it worked when nine out of ten drafts go out with only minor edits.
That third blank is the one people skip, and it’s the one that matters. If you can’t describe success in concrete terms, you don’t have an agent yet. You have a chatbot with extra steps.
The Four Parts Every Agent Needs
Every working agent I’ve built or debugged has the same four parts, whether it’s thirty lines of Python or a drag-and-drop workflow:
- A goal narrow enough to test. “Handle customer support” is not a goal. “Draft refund replies for orders under 60 days old” is.
- Two to four tools. Past four, models start reaching for the wrong one at the wrong moment.
- Memory of what it has already done. Even a plain running list of completed steps eliminates most duplicate-action bugs.
- A stop condition. A hard cap on steps plus a clear success signal. Without both, loops happen. Every time.
A Concrete Build: Refund Triage in Four Steps
1. Write the job as a system prompt
Give the agent its role, its boundary, and its output format in plain language: “You handle refund requests for an online store. Look up the order before replying. Refunds are allowed within 60 days of purchase. Draft a reply in the customer’s tone and never send it. If the order is older than 60 days, draft a reply that explains the policy and flag the thread as needs-human.”
Notice what isn’t there: no sprawling list of edge cases, no stern warnings about accuracy. Boundaries beat pep talks.
2. Hand it the smallest set of tools that works
Four functions cover the whole job: one to fetch an order, one to check policy against a date and amount, one to save a draft, and one to flag a thread. Each does a single thing and returns something readable. If a tool spits out raw JSON with forty fields, wrap it so the agent sees the five fields it actually needs. You’re designing for a reader, not a database.
3. Add the loop and the brakes
Agents work by repeating: think, call a tool, read the result, repeat. That loop needs rails. I cap mine at six steps, log every one, and require the agent to stop and ask for a human when a lookup fails twice. If you want a deeper look at how frameworks keep that decision-making predictable instead of chaotic, this write-up on keeping an agent’s decisions on the rails is worth twenty minutes.
4. Break it on purpose
Feed it the ugly inputs before your customers get the chance. An email with no order number. A refund request for something never purchased. An all-caps message from a customer on their fifth complaint. A perfectly normal request that should really be a one-step task.
If your agent invents an order number just to keep moving, you’ve found the bug that matters. This walkthrough of an n8n agent that triages support tickets hits the same wall from a different direction, and the fixes transfer directly.
When You’d Rather Skip the Code
Not everyone needs a Python scaffold, and you shouldn’t feel bad about skipping it. The same four parts assemble fine in visual tools. This Zapier AI agent built step by step, with two worked examples, shows how far you can get with a trigger, a couple of actions, and a short prompt. And if you’ve never written a line of code, there’s a solid guide to making AI agents without engineering experience that covers the vocabulary and the tool trade-offs.
The trade-off is real, though. No-code gets you to 80% in an afternoon and starts fighting you around step twelve of any complicated reasoning chain. Code is slower to start and far easier to debug later.
The Mistakes That Show Up in Almost Every First Build
I’ve reviewed enough of these to spot the pattern. The failures are rarely about the model.
- Success criteria so vague that nobody can say whether the agent passed or failed.
- Six tools when three would do, which produces hesitating, indecisive behaviour.
- No logging, so the first bug you hit is invisible.
- Testing only the happy path, then shipping to real users who don’t read instructions.
- Letting the agent send email or move money without a human approval step in version one.
Measuring Whether It Actually Works
Pull twenty real examples from your inbox and run them twice. For each one, log three things: did it finish without a human rescue, did the output need editing, and how many steps did it take. Two runs matters because a flaky agent that works on alternating attempts is far more expensive than one that fails cleanly.
My rough bar for shipping: 85% of drafts usable with light edits, under eight steps on average, and a few cents of token cost per run. If you’re nowhere near those numbers, the problem is almost always in the tools or the goal, not the model.
Ship the Ugly Version First
The first refund agent I built was forty lines and a CSV of logs. In its first week it drafted twelve replies, five of which a human sent with a single edit. That’s a win, and it taught me more than three weeks of theorising would have.
Add tools one at a time after that, and only when a real failure demands it. If a chat-style interface would get your team using it faster, an approach like Poke, which turns agent building into a text-message conversation, is worth a look. Start with the smallest version that can finish one job end to end, watch it fail, and fix the failure you can actually see. That loop, repeated a few times, is how you make AI agents that people keep.

