Six months ago a 42-person logistics company had one genuinely good AI project. It read inbound carrier emails and drafted replies for the three people on the dispatch desk. It saved them roughly nine hours a week. Then the CEO asked the obvious question, can we do this everywhere, and everything got harder.
That gap between one working AI project and an organisation that actually runs on it is where most scale AI efforts quietly stall. Not because the models fall over. Because nothing around them was built to travel. The prompts live in one person’s notes, costs stay invisible until the invoice lands, and every new team rebuilds the same thing from scratch.
What follows is the sequence I use, in the order that matters.
Step 1: Pick the workflow before you pick the model
The instinct is to start with a model comparison. Resist it. Start with a workflow and hold it to three tests:
- It happens at least 20 times a day, because volume is what creates the payoff.
- A human touches every instance today, so you can measure the before.
- Someone already double-checks the output, so you aren’t inventing a review process from nothing.
Carrier email triage passed all three. “Build us an AI strategy” failed instantly. No volume, no baseline, no way to tell whether it worked.
Then write the current state down in numbers rather than adjectives. For that dispatch team: 340 emails a day, four minutes average to triage, 22% routed to the wrong depot, two escalations a week that reached a customer. That single page becomes your baseline and your business case, and you will come back to it in week twelve.
Step 2: Attach a dollar figure to the line you’re moving
340 emails at four minutes each is 22.7 hours of work a day. Loaded at $28 an hour, that’s about $635 a day, or roughly $160,000 a year. A 60% cut in triage time is worth close to $96,000 annually.
Now set a ceiling. Cap total spend, including licences, inference and your own engineering hours, at 30% of the value you expect to capture. Here that’s about $29,000 a year. Any quote above the ceiling is an automatic no, and that one rule ends more internal arguments than any benchmark you could run.
Inference is the line item people guess wrong. Modelling it properly rather than treating it as a rounding error changes how you build. If you’re weighing hosted models against running your own, a clear-eyed look at Azure OpenAI’s costs and where it genuinely fits shows the shape your spreadsheet should take.
Step 3: Build the plumbing nobody volunteers for
Before use case number two, five things need to exist. None of them are exciting. All of them are why the second use case takes three weeks instead of three months.
- A prompt registry, versioned like code, so you can see what changed and roll back.
- Structured logging of every request and response with model, tokens, latency and cost attached.
- An evaluation set of 100 to 200 real examples with expected outputs.
- One access story: which teams may send which data to which provider.
- A kill switch that any team lead can hit without a meeting.
Six weeks of clean logging tells you which model is actually worth paying for, which prompts are quietly failing, and which team is burning budget on a workflow nobody uses. If you’d rather not design this from zero, the approach in setting up an AI assistant for real work maps onto internal tooling just as well as it does onto a single assistant.
Step 4: Give every team the same rails
Teams that build alone produce twelve slightly different versions of the same thing. Give them one set of rails instead.
A shared request template
Context first, then task, then the constraints, then two worked examples. Anyone writing a prompt starts here. It cuts the “why is our output worse than theirs” conversation to nothing.
A library of golden examples
Fifteen of the best real inputs and outputs per use case, curated by the owning team. New joiners read them. Evals run against them. They drift, so someone reviews quarterly.
A routing rule, not a favourite model
Default to the cheap fast model, escalate to something heavier only when a confidence trigger fires. Model line-ups shift fast, so revisit the routing once a quarter rather than rebuilding per project. This run-through of the current model landscape and what each tier is good for is a reasonable way to keep that rule current.
Step 5: Automate the review loop, not just the work
Scaling AI isn’t about removing humans. It’s about choosing exactly where the humans sit.
For the dispatch tool, anything scored under 80% confidence got flagged for a human before sending. At launch that was 12% of drafts. By week eight it was 5%. Ten minutes of a supervisor’s morning, every morning, and nobody’s reputation on the line.
Alongside that, one person reviewed 20 randomly sampled outputs each Monday and logged corrections. The correction rate went from 14% to 4% over seven weeks, and that number, not the accuracy score on a launch demo, became the metric the leadership team actually asked about.
The feedback structure here is the same one that makes a conversational AI assistant hold up under real traffic. Sample, score, correct, re-run the eval set. Repeat weekly.
Step 6: Budget, govern and expand, in that order
Track cost per task from day one. For triage that was $0.11 per email once logging matured. It’s the number that tells you whether use case five is safe to green-light, because it lets you forecast before you commit.
Governance fits on one page: what data may go where, who approves a new use case, and what happens when the model is wrong. Anything longer gets ignored.
Then set an expansion gate. Don’t start use case four until two and three have run 30 days without a human rescuing them. The habits that carry an AI pilot into production are mostly this: patience about sequencing, and a refusal to add surface area while something is still wobbling.
The asset that decides whether any of this holds
None of the six steps survive without one named owner. Not a steering committee, not an “AI working group” that meets monthly. One person with roughly 20% of their week protected for it.
Their job is narrow and unglamorous. Keep the prompt registry honest. Refresh the eval set. Make costs visible to the finance team. Enforce the expansion gate when a director is impatient. Answer the question “can we point this at the contracts folder” with a policy rather than a shrug.
Companies that scale AI well usually look boring from the outside. Same template everywhere, one dashboard, a weekly sampling ritual, a spend ceiling nobody argues about because it was agreed in month one. The interesting part, the part that gets written about, is about 10% of the work. The other 90% is the owner, sitting down on a Monday, reading twenty outputs and correcting four of them.

