A support inbox with 200 unread tickets. A competitor’s pricing page that someone checks by hand every Monday morning. Forty PDFs that need summarising before a Friday board meeting. These are the jobs where AI agents stop being a demo and start being useful, and none of them are glamorous. They’re just repetitive, tedious, and rule-shaped enough that a loop with a few tools can handle them.
Most people get stuck at the wrong end of the problem. They read about agent frameworks, pick one, then go looking for something to build. The order is backwards. Here’s the sequence that actually works, with a worked example running through it.
Step 1: Pick a job with a clear finish line
An agent needs to know when it’s done. “Answer all customer questions” has no finish line and a huge blast radius. “For every new support email, pick one of three labels and draft a reply using the refund policy doc” does. Same technology, wildly different outcomes.
Three tests before you write a line of code:
- Repeatable. The task happens at least weekly. One-off jobs are faster to do by hand.
- Verifiable. You can check the output in under a minute. If reviewing the work takes longer than doing it, you’ve picked wrong.
- Bounded. A mistake costs an apology, not a lawsuit. Start where errors are cheap.
Write the job down as a single sentence with a trigger and an output. “Every weekday at 7am, check this pricing page and post to Slack only if something changed” is a spec. “Help with competitor research” is a wish.
Step 2: Build the loop before anything else
Strip away the frameworks and an agent is four lines of logic. Send the model a message plus a list of available tools. If it asks to call one, call it, paste the result back into the conversation, and loop. If it answers in plain text, you’re finished. Everything you’ve read about planning, reflection, and multi-agent swarms is scaffolding built around that core.
Get the bare loop working first. My first useful agent did exactly one thing: every weekday it fetched a competitor’s public pricing page, compared it against a snapshot saved as JSON from the previous day, and posted a Slack message only when a number had moved. Nineteen lines, no framework, no vector database. It ran for eight months.
That version taught me more than any tutorial, mostly because when it broke I could see exactly where.
Step 3: Tools are where an agent earns its keep
An agent that can only talk is a chatbot with a loop. The moment it can search your documentation, look up an order, or send a message, it can finish work instead of describing it.
Wiring up a dozen SaaS integrations by hand is its own multi-week project, which is why Composio packages hundreds of tool integrations behind a single interface. Worth knowing about before you start writing OAuth flows for Gmail and HubSpot yourself.
A sensible starting toolset for a support triage agent:
search_policy_docs— keyword or vector search over your help centre articlesget_order_status— one read-only API call, nothing moredraft_reply— writes to a draft field, never sendsapply_label— sets one of three tags
Keep the list under about eight tools. Every addition is another thing the model can misuse, and every tool description eats prompt space that could be going to instructions.
When the job only exists in a browser
Plenty of tasks live behind a login and a badly designed form with no API in sight. That’s a different class of problem, and it’s where browser agents that click, type and navigate on your behalf come in. They’re slower and flakier than API calls, so treat them as the fallback, not the default.
Step 4: Give it memory, or it will ask the same question every morning
A stateless agent starts from zero on every run. For a one-shot digest that’s fine. For anything conversational, it’s fatal. You end up re-feeding context you already had.
There are three kinds of memory worth separating. Short-term is the current conversation. Long-term is durable facts: this customer is on the enterprise plan, they always want invoices in euros. Episodic is what happened last run, which is how an agent avoids repeating a failed approach.
Start cheap. A single JSON file holding yesterday’s state solves most of it. Reach for something heavier when you need search across thousands of past interactions, and Letta treats agent memory as a first-class part of the architecture rather than an afterthought bolted on later.
Step 5: Constrain outputs with types, not vibes
Ask a model for “a summary and a category” and you’ll get clean JSON on Monday and a bulleted list with a friendly preamble on Tuesday. Anything downstream that parses that output breaks.
Define the shape you want. A refund decision should be an enum with three values, not free text. A price should be a number with a currency code attached. Frameworks such as PydanticAI bring type validation to agent outputs, so a malformed response triggers a retry instead of a silent failure three steps later.
The validation layer is also your cheapest debugging tool. When a run goes wrong, the first question is always “what did the model actually return?”
Step 6: Decide the blast radius before you walk away
The failure that ends projects isn’t a bad answer. It’s an agent with more access than it needs, running unattended. Files uploaded to a public bucket, credentials leaking into logs, data sent somewhere it shouldn’t go. It has already happened at least once, when unsecured agents published 53 user images without anyone at the lab realising.
Four rules that cover most of the risk:
- Read-only by default. Write access is granted per tool, deliberately.
- Scoped credentials. The agent’s Gmail token should reach one label, not the whole mailbox.
- Human approval for anything irreversible. Refunds over $50, emails to more than one recipient, deleting rows.
- Log every tool call with its inputs and outputs. When something odd happens at 3am, this is the only record you’ll have.
Step 7: Test against a graded set, not a feeling
Agent testing looks different from software testing because the same input can take different paths on different runs. You can’t assert one correct output.
Instead, collect twenty real examples. For each, write down what you’d accept as a good answer. Then grade every run against that bar: pass or fail, judged by you or by a second model. A pass rate you can track beats a gut feeling every time.
Watch two numbers alongside the pass rate. Total steps, because an agent that finishes in four calls is cheaper and easier to debug than one that wanders for thirty. And cost per run, because a task that takes $2 of tokens to complete might be cheaper to do by hand. Cap iterations hard. Thirty loops means something is stuck.
The whole thing, end to end
Back to that support inbox. The finished spec looks like this: new email arrives, agent checks order status and searches policy docs, picks one of three labels, drafts a reply, and stops. Memory holds the last fifty interactions with that customer so it doesn’t ask for an order number twice. Refunds above $50 get flagged rather than drafted. Every tool call is logged.
The eval set is twenty historical tickets with the label a human gave them and the outcome that actually happened.
In the first fortnight it handled 61% of tickets without a human touching them. Almost all the misses were angry customers, where the draft was perfectly fine but the tone was wrong, plus a handful of edge cases where the label was off. That’s a normal result. The wins cluster in the boring middle, and the edges stay with people.
A 45-minute version you can run today
If none of that feels concrete enough yet, build this instead. Take ten URLs from a conference speaker list and have an agent fetch each one and return a table with company name, pricing tier, and a one-line positioning statement.
One tool: fetch_url. No memory. No writes. Twenty minutes of setup and twenty-five of fiddling with the output format until the table stops drifting. When it works, check three rows by hand and see how often the model quietly invented a pricing tier.
That exercise gives you the loop, a real tool, and a reason to care about validation. From there, every addition is a constraint: memory so it stops repeating itself, types so the output holds shape, permissions so a bad day stays recoverable. Layer them one at a time and the boring job eventually runs itself.

