A support team I worked with shipped a refund agent that handled 40% of their ticket volume for three weeks without a single complaint. Then a model provider rotated a snapshot on a Tuesday. By Friday the agent was issuing partial refunds to customers who had only asked about shipping delays. Nobody caught it for two days, because the only thing anyone was watching was the final answer text, and the final answer text looked fine.
The gap between “the agent produces plausible output” and “we know whether the agent is doing its job” is what AgentOps fills. What follows is the sequence I now use when putting an agent into production, in the order it actually matters. Skip a step and you will usually pay for it later, often at 2am.
Step 1: Write down what “working” means, in numbers
Before adding a single log line, get the team to agree on three to five measurable outcomes. Not vibes. Numbers with thresholds attached.
An example service contract for a customer support agent
- Resolves at least 55% of tier-1 tickets with no human edit
- Escalates to a human when its own confidence score drops below 0.7, targeting a 20-30% escalation rate
- Costs under $0.25 per conversation, all retries included
- p95 end-to-end latency under 9 seconds
- Zero policy citations that do not exist in the knowledge base
That last one matters more than the others. An agent that invents a refund policy is worse than an agent that says nothing. Numbers like these give you something to alert on, something to regress against, and something to show a skeptical VP when they ask why the agent is only handling half the queue.
Step 2: Instrument the whole trajectory, not the last message
Most teams start by logging the user input and the final response. That tells you nothing about why the agent went sideways. You need spans: one for each retrieval, each tool call, each model invocation, nested under a single run ID that you propagate through the entire agent loop.
This is where a proper tracing setup earns its keep. If you are choosing tooling, the practical guide to tracing and evaluating LLM apps with LangSmith walks through the span structure and what each field buys you at debug time.
Fields to capture on every span
- Prompt template hash and version identifier
- Model name plus the exact snapshot, not just “gpt-4-class”
- Retrieved chunk IDs and their similarity scores
- Tool name, arguments, latency, and any error payload
- Token counts in and out, per span
- The user-visible output, stored exactly as rendered
Nine times out of ten, when an agent misbehaves, the answer is sitting in the tool-call sequence. It retrieved the wrong document because the query rewrite dropped a product name. It called the order lookup twice because the first response was a timeout it never checked. You cannot see any of that from the final message.
Step 3: Build your eval suite out of real failures
Do not start with synthetic test cases. Start with production traces. Wait until you have roughly 200 real conversations, then sit down and label the failures by hand. It takes an afternoon. Every incident you find becomes one permanent eval case, which means your regression suite grows exactly as fast as your agent’s real-world weaknesses.
Split the suite into two tiers. Deterministic checks handle the boring stuff: does the tool call parse, is the order ID the right format, did the agent call the refund API before confirming eligibility. LLM-as-judge handles tone and faithfulness, and it should always be graded against a rubric with a written rationale, not a bare numeric score.
A 120-case suite that runs in four minutes is plenty to start. Run it on every prompt change, and block the deploy if the pass rate drops more than three points.
Step 4: Give every run a budget and enforce it at runtime
Cost and latency are agent behaviours, not infrastructure metrics. An agent that averages 14 tool calls when it needs 4 is a bug, not a cost problem.
Set hard ceilings per run: eight tool calls maximum, 12,000 tokens, 15 seconds wall clock. When a run hits the ceiling, kill it and fall back to a scripted response plus a human handoff. I watched a research agent go from $0.61 per query to $0.28 simply by capping tool calls at six and instructing it to answer with what it had. Judged quality dropped 1.4 points out of 100. That trade was obviously worth making, but nobody would have made it deliberately without the number in front of them.
Step 5: Build the escape hatch before you polish the happy path
Teams love tuning the successful path. Escalation design gets left until launch week, and then it gets bolted on badly. Do it early, because your escalation triggers double as a safety net for everything else that goes wrong.
Triggers that should always hand off to a human
- Top retrieval similarity below your threshold, meaning the knowledge base simply does not cover the question
- Two consecutive tool errors of any kind
- The user repeats the same question twice in a row
- Any request involving money above a set limit, account deletion, or legal language
Write the handoff summary as a fixed template: what the customer wants, what the agent tried, what it found, and the single question the human needs to answer. If a rep cannot pick up the conversation in 30 seconds, your escalation is just a different kind of failure.
Step 6: Alert on behaviour, because nothing else will catch a silent model change
Generic API monitoring will tell you the endpoint returned 200. It will not tell you the agent got dumber. The reason is structural, and the breakdown of why agent monitoring breaks the assumptions of MLOps tooling is worth reading before you wire up dashboards, because a lot of the standard machinery simply does not transfer.
What works instead: sample a slice of live traffic, say 5%, and run your eval judge over it nightly. Track the pass rate week over week. Alert if it drops five points. Track the distribution of tool-call counts and escalation rates the same way. A model snapshot rotation, a retrieval index rebuild, or a knowledge base edit will show up as a shift in one of those distributions long before a customer complains.
If you are standardising on a framework, the NVIDIA NeMo Agent Toolkit ships profiling hooks that give you tool-call timing and token accounting for free. Worth knowing where those hooks stop short of what you actually need, though, since framework telemetry rarely covers eval scores or semantic drift.
Step 7: Treat prompt edits as deploys
A prompt change is a code change that skips code review. Version every prompt, tag it in your tracing tool, and run it through the regression suite before it touches traffic. Then use shadow mode: send the new prompt 100% of live requests, discard the output, and compare it against what the live agent produced. Two days of shadow comparison catches things a 120-case suite never will.
After that, a 5% canary with the alerts from step 6 armed. Ramp to 25%, then 100% over about a week. Keep the previous prompt ready to restore in one command, because the version that worked in March may quietly fail in June once the product catalogue has changed underneath it.
The order teams usually get wrong
Almost every failed rollout I have seen followed the same shape: heavy investment in the prompt and the tools, light investment in metrics, tracing, and escalation. The result is an agent that demos beautifully and then sits at 15% adoption because nobody trusts it and nobody can explain its failures.
The operational playbook for running agents in production makes the same point from the other direction: the engineering that gets an agent running is a fraction of the work, and the rest is the unglamorous loop of measuring, regressing, and shipping careful increments. Build the measurement layer first, even when it feels like a detour. It is the only thing that lets you move fast later without guessing.

