Six hundred tickets a week, three support engineers, and a tagging spreadsheet nobody updates. That was the setup a friend described to me before she started building a triage agent with Google Vertex AI Agent Builder. Two weeks later, the agent reads every incoming ticket, classifies intent, pulls the matching help-centre article, drafts a reply, and routes anything angry or billing-related to a human. It handles roughly 70% of first-touch work.
This is the walkthrough I wish she’d had on day one. We’ll build that same triage agent from scratch, step by step, and I’ll flag the parts that quietly ate entire afternoons.
What Agent Builder actually gives you
Strip away the product pages and there are four moving pieces you’ll spend your time on:
- Agent instructions — the system prompt that defines role, tone, and hard limits.
- Tools — OpenAPI functions, Cloud Functions, or prebuilt connectors the agent can call.
- Data stores — retrieval over your own docs, wired to Vertex AI Search.
- Sessions and examples — conversation state plus few-shot samples that teach behaviour.
That’s it. Most failures I’ve seen trace back to weak instructions or a broken tool schema, not the orchestrator. If you want the conceptual picture before you open the console, this practical guide to building agents that actually work is a good place to get your bearings.
Step 1: Write the instructions like a job description
Vague prompts produce vague agents. Here’s a trimmed version of what we shipped:
You are a support triage agent for Acme Analytics, a B2B dashboard product. For each incoming ticket: classify intent into one of [billing, bug, how-to, cancellation, other]. If intent is billing or cancellation, escalate immediately without drafting a reply. Otherwise search the help centre and draft a reply under 120 words. Never invent pricing, refund terms, or SLA numbers. If retrieval returns nothing relevant, say so and escalate.
Two things matter there. The intent list is closed, so the model can’t invent new categories. And the escalation rule is stated before the drafting rule, which measurably reduced the number of refund guesses.
Give the agent permission to refuse
Refusal rules are the cheapest safety net you’ll ever write. Add a line for every category you’re nervous about: legal threats, data deletion requests, someone who’s already been escalated twice. Without them, a helpful model will happily improvise a GDPR response.
Step 2: Hand it some tools
Tools are where the agent stops talking and starts doing. We registered four:
get_customer— pulls plan tier, MRR, and account age from BigQuery.search_help_center— queries the Vertex AI Search data store.add_ticket_note— writes the classification back to Zendesk.escalate_to_human— assigns to a queue with a reason code.
Keep tool descriptions brutally clear about inputs and outputs. If a parameter is optional, say what happens when it’s missing. Half the debugging time in our first week went to a tool that silently returned an empty list instead of an error, so the agent confidently told customers their account probably didn’t exist.
Ground it on your own documentation
Point search_help_center at a Vertex AI Search data store built from your help centre export. Chunk size is the dial that matters most. At 1,000 tokens we got mushy answers that blended three articles together; at 400 tokens, groundedness scores jumped noticeably. The ingestion and retrieval patterns behind this are the same ones LlamaIndex was built around, and reading up on how it handles chunking and retrieval will save you from re-deriving it yourself.
Step 3: Use flows for anything with a fixed order
Agents are excellent at deciding. They are unreliable at remembering to always do step three. Refund processing, password resets, and data export requests all have a required sequence, so wrap each one in a flow with deterministic steps and expose the flow to the agent as a single tool.
The agent picks the flow; the flow guarantees the order. If you’ve built anything in a visual builder such as Flowise, where nodes pass state along edges, this will feel immediately familiar, and you can prototype the flow logic there before rebuilding it natively.
Step 4: Evaluate with real tickets, not vibes
Vertex’s evaluation tooling lets you upload a golden set. Ours was 80 real tickets with expected intent labels and escalation decisions, pulled from the previous quarter and scrubbed of customer names. Four metrics ran on every instruction change: intent accuracy, tool-call correctness, groundedness, and safety.
Concrete before-and-after numbers from our tuning week: groundedness went from 0.71 to 0.93 after shrinking chunks and adding a citation requirement. Intent accuracy moved from 0.88 to 0.91. Tool-call correctness was the slow one, and it only improved once we rewrote two tool descriptions in plain language.
Build an adversarial set too. Ours included an empty ticket body, a Spanish message, a pasted 4,000-character stack trace, a reply misfiled from an unrelated thread, and one all-caps refund demand. The last one broke the agent until we added the escalation-first rule.
Step 5: Ship it in draft mode, then widen the gate
Deploy to Agent Engine with the agent drafting replies that a human approves. Track approval rate per intent category. When a category holds above 90% approval for two consecutive weeks, switch it to auto-send. Ours took eleven days to clear how-to questions and about five weeks for bugs, which is reasonable given how often bug reports are really billing complaints in disguise.
Cost is lower than people expect. A triage workload of 600 tickets, each with three or four tool calls, lands somewhere in the $30–60/month range on a Flash-class model. Swapping to a Pro-class model for everything multiplied that by roughly nine, and the only category that measurably benefited was billing disputes.
Log full sessions from day one. In week one you’ll find the agent inventing a 30-day refund window that doesn’t exist. That isn’t a model problem, it’s a missing refusal rule, and the session log is what tells you which one.
When a managed platform isn’t the right call
Agent Builder earns its keep when you want Google’s retrieval, evaluation, and deployment plumbing handled for you. If your logic is a tangle of custom chains, you need everything running inside your own VPC, or you’re already deep in a Python codebase, a library-first route is often faster to reason about, even though you own more of the machinery. It’s worth understanding when LangChain’s approach makes more sense before you commit either way, because migrating an agent mid-project is genuinely unpleasant.
Your second week looks different from your first
Week one is plumbing: instructions, tools, a golden set, a deployment. Week two is where the agent starts earning trust. Add a category to auto-send, then check the approval rate two days later. Pull the ten lowest-groundedness sessions and fix whichever document is missing or stale. Watch for tool calls that return empty results, since those are usually schema bugs rather than model failures.
Keep a running list of every escalation reason code and review it weekly. When one reason dominates, that’s your next flow, not your next prompt tweak.

