Your support queue takes about 900 tickets a week and four people tag them by hand. Nobody’s job title says “tag tickets,” which is exactly why it keeps sliding to the bottom of the pile. This is the kind of small, repetitive decision that Google Cloud AI is genuinely good at, and it’s a good first project because you can actually measure whether it worked.
What follows is the walkthrough I’d hand a team starting from zero. One real use case, eight steps, and the numbers that keep it from turning into an expensive science experiment.
Step 1: Pick One Decision, Not a Platform
The most common way these projects die is a kickoff meeting that ends with “we should adopt AI.” That’s not a project, it’s a mood. Pick a single decision the model makes, hundreds of times a week, where a wrong answer is cheap to catch.
For the ticket example, the decision is: route this ticket into one of four buckets (billing, damaged bike, refund request, general). Success looks like 85% agreement with your human labels, with the other 15% going to a person anyway. That’s a defensible target. “Improve customer experience” is not.
If you’re still deciding which of Google’s services to reach for, it’s worth twenty minutes with a plain-English overview of Google Cloud AI, explained without the marketing, because picking the wrong product here costs you weeks.
Step 2: Set Up the Project So You Don’t Regret It in March
Four things, in this order:
- Create separate projects for dev and production. Not folders, not labels. Projects. It’s the only clean way to keep an experiment from touching live billing.
- Enable only the APIs you need: the Vertex AI API, plus Document AI or Speech-to-Text if your inputs aren’t already text.
- Make a service account with the narrowest role that works (Vertex AI User is usually enough) and download the key once.
- Set a budget alert at a number that would annoy you, not a number that would ruin you. For a pilot like this, $200 is generous.
Pick your region and stay in it. Model availability, latency, and data-residency rules all vary between us-central1 and europe-west4, and moving later means re-testing everything.
Step 3: Prototype the Prompt in Vertex AI Studio Before Writing Any Code
The console’s prompt playground is free to poke at and it’s where you should spend an entire afternoon. Not in your editor. The turnaround is seconds instead of minutes, and non-engineers on your team can read the output and argue with it.
Hand the model your taxonomy, verbatim
List all four categories with one real example ticket under each. Don’t paraphrase your help-centre definitions, paste them. Models follow concrete examples far better than descriptions of examples.
Then feed it the worst tickets you have
All-caps rage, three languages mixed in one sentence, a photo of a cracked frame pasted with no words, and the inevitable “I want to speak to a manager.” If your prompt survives ten of those, it’s ready to move.
Start on a Flash-tier Gemini model and see how far it gets. It’s fast and cheap enough that you can run the whole queue nightly without thinking about it, and the team behind it is the same DeepMind research group that built Gemini in the first place. Only reach for a larger Pro model if you can point at specific tickets where Flash gets it wrong.
Step 4: Move the Prompt Into Code (It’s About Thirty Lines)
Use the Vertex AI SDK for Python. You send the ticket text plus your system instructions to generate_content, but the important part is the response schema: ask for JSON with a category field and a confidence field, and set the response MIME type to application/json. Now you get parseable output every time instead of prose you have to scrape.
Two settings matter here. Temperature at 0.1, because you want the same ticket to land in the same bucket every time. And a batch path, because if nightly routing is fine, you can send 900 tickets in a handful of calls rather than 900 separate ones.
Worth knowing: this pattern isn’t exclusive to Google. If your team already lives in Microsoft’s stack, the same prompt-and-schema approach maps across to Azure OpenAI Service almost one for one, which is a decent hedge against getting locked in.
Step 5: Handle the Inputs That Aren’t Clean Text
Real tickets arrive as voicemails, scans, and photos. Each of those has a specific tool, and each one converts the mess into text that feeds your existing prompt. Don’t build a second pipeline.
- Document AI for PDFs and scanned receipts. The form parser runs roughly six cents a page, so only route the tickets that actually contain an attachment.
- Speech-to-Text for voicemails, at about 1.6 cents a minute. A few hundred calls a month lands around $5.
- Cloud Vision for damage photos, using label detection to flag whether a bike image even needs a human look.
Chain them so the extracted text hits the same Gemini prompt as everything else. One prompt to maintain, one evaluation set to re-run.
Step 6: Build a 200-Ticket Golden Set Before You Ship
This step is the one teams skip, and it’s the one that decides whether the project survives contact with your support lead.
Pull 200 real tickets from the last quarter. Have two people label them independently, then sit down and resolve every disagreement. Those resolved disagreements are where you’ll discover your categories overlap more than you thought.
Run the model against all 200 and compare. Expect 80-90% agreement on a clean taxonomy. If you land at 65%, look at the failures in groups rather than one by one, because if 40% of the misses are “billing versus refund,” your taxonomy is the problem, not the model. Tighten the definitions and re-run.
Keep this set. Every time you touch the prompt, run it again. It’s a regression test that takes two minutes and has saved more launches than any code review I’ve seen.
Step 7: Deploy on Cloud Run and Do the Cost Math
Cloud Run is the right home for this. Set minimum instances to zero so idle time costs nothing, and let concurrency sit around 20. Your service only wakes when a ticket arrives.
Now the numbers, because this is where people panic unnecessarily. At 900 tickets a day, an average ticket is roughly 400 input tokens (prompt plus message) and 60 output tokens (the JSON). That’s about 0.36 million input tokens and 0.054 million output tokens per day. On Flash-tier pricing, call it ten cents per million in and forty cents per million out. Total: under two dollars a month for inference.
The bill rarely comes from the model. It comes from Document AI pages, verbose logging kept at full retention, and one colleague who set minimum instances to 1 “just for testing” and never changed it back. If you want a clearer picture of where AI infrastructure spending actually accumulates, what you’re really renting and what it costs is a useful reality check before your first invoice.
Step 8: Keep a Human in the Loop and Watch for Drift
Route anything below 0.8 confidence straight to the human queue, untouched. That single rule buys you most of the safety you need, because the model is generally well-calibrated about when it’s guessing.
Log every prompt, response, and final human correction to BigQuery. Then spend twenty minutes a week reading the misses. Not the individual ones, the patterns. You’ll spot things like a new product name the prompt has never seen, or a promotions campaign that floods one category for a fortnight.
Drift is slow and boring. Vocabulary shifts, accuracy slides from 88% to 81% over three months, and nobody notices because the system still mostly works. Re-run the golden set monthly and you’ll catch it while it’s still cheap to fix.
When that monthly review comes back clean two months running, don’t add a dashboard. Move on to the next decision worth automating, and leave this one running quietly in the background where it belongs.

