Most teams don’t have an AI problem. They have a shortlisting problem. There are hundreds of tools on the market, half of them do the same three things, and every free trial expires long before anyone has pointed it at real work.
The fix isn’t a longer comparison spreadsheet. It’s a short, boring process you can finish inside a week. Here’s the one I’d run if you gave me five working days and a stubborn operations manager.
Step 1: Name the job before you name the tool
Write one sentence that describes a task, not an ambition.
Weak: “Improve productivity across the team.”
Useful: “Turn a supplier’s PDF invoice into a row in our purchase ledger in under four minutes.”
The second version tells you what to test, who owns it, and whether it worked. It also tells you what a bad result looks like: a row with the wrong VAT figure sitting in it.
Pick exactly one job for the week. A 12-person accountancy practice that did this chose “reconcile bank statement lines against the client ledger” rather than “use AI for accounts”. Narrowing it down cut their evaluation from roughly three weeks to two afternoons.
The sentence template
[Verb] a [specific input] into [specific output] in under [time limit]. If you can’t fill in the blanks with something a colleague could check, you’re not ready to shop.
Step 2: Build a shortlist of three, not thirty
Plenty of round-ups you’ll find are paid placements, so start with a curated list of tools that have already survived contact with real work. This piece on artificial intelligence tools that are actually worth your time is a reasonable first filter. Then cut it to three using your own constraints:
- Where the work already lives. If invoices arrive in Gmail, a tool that only accepts manual uploads will be abandoned by Thursday.
- Data handling. Client names, health data, anything under NDA needs a clear answer on retention and model training. Get it in writing.
- Price at your real volume. A $20 monthly plan that caps at 100 runs is a $400 monthly plan in disguise.
- Escape route. Can you export your work if you leave? If not, that’s a one-way door.
Three tools. Not five. Decision research keeps landing in the same place: past four or five options, people stop comparing and start guessing.
Step 3: Test on your own messy work
Demo data is clean, polite and useless. Pull ten real examples, including the ugly ones. The scanned receipt that’s slightly rotated. The email thread where three people reply at once. The invoice from the supplier still using a 2014 template.
Run all three tools against those same ten examples and score four things:
- Time to a usable first draft
- Minutes you spent fixing it (the number everyone forgets)
- What went wrong, specifically
- Cost per completed task
Do it properly and it becomes a real experiment rather than a vibe check. There’s a solid walkthrough of how to run an artificial intelligence test that actually tells you something, and it’s worth reading first, because the classic mistake is testing twenty easy cases and declaring victory.
Watch out for leaderboard thinking. A tool that tops a public ranking can still choke on your scanned receipts. The reasoning behind why benchmarks alone don’t tell you much about real-world performance applies directly here. Your data is the only benchmark that matters.
What a decent scorecard looks like
Tool A: 4 of 10 usable, average 11 minutes of editing. Tool B: 9 of 10 usable, average 2 minutes. Tool C: 9 of 10 usable, 90 seconds of editing, but it quietly dropped a line item on two invoices. Being wrong quietly is worse than being wrong loudly, and Tool C loses on that alone.
Step 4: Write the context you’ll reuse
Tools get noticeably better once you stop typing fresh instructions every single time. Spend one hour building a small context pack:
- A 150-word note on how your team writes: tone, terminology, words to avoid
- Three examples of a finished output you’d happily send
- The five mistakes the tool made in testing, rewritten as instructions
Save the whole thing as one reusable prompt and park it somewhere the team can actually find it. The same logic applies to finding information rather than drafting it. This workflow for getting reliable answers from an AI search engine follows an identical shape: frame the question well, check the sources, repeat.
Step 5: Give the tool a home in the workflow
This is where most rollouts die. The tool works in isolation and then never gets opened, because it sits outside the place where the work happens. Don’t ask people to open a new tab and paste things in. Put the tool where the bottleneck is.
If invoices land in a shared inbox, the extraction step belongs in that inbox. If reports start life as a spreadsheet, start there. A step-by-step guide on putting an artificial intelligence platform to work for a real team covers this in more detail, including how to hand over ownership so it doesn’t quietly become one person’s side project.
Step 6: Measure, then decide
Pick three numbers before you start and check them at the end of week two:
- Time per task, before and after
- Review burden, meaning how many minutes a human still spends checking
- Error rate, counted honestly, including the ones you caught yourself
A realistic result on a repetitive task is a 60 to 80% cut in handling time with the error rate unchanged. “It feels faster” is not a result.
Set a kill criterion in advance. Something like: if a human still spends more than half the original time reviewing the output, drop it and go back to Tool B. Teams that write this down abandon bad picks far more quickly than teams that don’t.
One more thing to watch
Think about what happens when the tool is wrong in a way nobody notices. A summariser that drops a clause. A categoriser that codes a legitimate expense as personal. Decide now who checks what, and how often, before the volume scales up and the checking quietly stops.
What week two looks like
Monday: hand the template to one colleague who wasn’t involved in testing, and watch them use it without any help. Tuesday and Wednesday: let them run it for real, and write down every question they ask. That list becomes your one-page guide. Thursday: fix the two most annoying problems and nothing else. Friday: compare the numbers against your week-one baseline, then either extend the trial to a second team or shut it down and move on.
One job, three tools, ten real examples, a reusable prompt, and a decision with a date attached to it. Not glamorous. It does beat another six months of comparing feature tables.

