A copywriter I know spent a Tuesday night doing something nobody should have to do: pasting her own work into an AI writing detector for the fourth time, screenshotting the score, and emailing it to an editor who didn’t believe she’d written it herself. The number that started the argument was 91%. The text was a 400-word product page built from two customer interviews and a factory visit.
She lost, mostly because you can’t negotiate with a percentage. You can, however, learn exactly what the tool in front of you responds to and how much weight its answers deserve. That takes about an hour of deliberate testing, and it turns a vague accusation into something measurable you can actually discuss.
Five checks, run in this order.
What the number on screen is really measuring
Most detectors combine two signals: perplexity, which asks how predictable each word is given the words before it, and burstiness, which measures how much sentence length and rhythm vary across a passage. A classifier trained on paired human and machine samples usually sits on top.
None of that means “91% of this text was written by a machine.” A high score means the prose matches a pattern the model associates with generated writing. Smooth, evenly paced, slightly generic text hits that pattern, which is why tax guidance, HR policy, and the work of anyone writing in a second language gets flagged so often. Detectors are imperfect instruments rather than judges, and treating the output as a verdict is where most of the damage happens.
Before you start: build a small control set
You can’t interpret an AI writing detector’s score without something to compare it against. Collect six plain-text samples of 250 to 300 words each. Length matters: plenty of tools truncate long inputs or pad short ones, and both distort the result.
- Two samples of your own unedited writing, one informal and one formal
- One sample written before 2022, if you can find any. This is your anchor.
- Two raw outputs from whatever model people keep accusing you of using
- One hybrid: model text you’ve rewritten by at least half, sentence by sentence
Keep the formatting boring. Strip headings, bold, and bullets from everything except the sample you’ll later use to test structure on purpose.
Check 1: score the pre-2022 sample first
Run the old sample before anything else. If an AI writing detector calls writing from 2016 “likely AI-generated,” the tool is broken for your purposes and every other result it produces is noise. You’ve just saved yourself an hour.
In my own runs, plain 2019-era business writing tends to land between 15% and 45% “AI” on the weaker tools and under 10% on the better ones. A reading above 60% on genuinely pre-LLM text says more about the detector than the writer.
Check 2: find the halfway point with staged edits
Take 250 words of raw model output and build four versions. Version zero stays untouched. Version one gets roughly 20% of its sentences replaced. Version two gets half rewritten. Version three is rebuilt from your own notes, keeping only the structure. Run all four, then write the scores down in a row.
Read the curve, not the number
What you want to find is the edit depth where the score falls below the tool’s own threshold, usually somewhere between 30% and 50%. If version three still scores 40% AI, you’re not looking at a writing problem. You’re looking at a measurement problem, and further editing won’t fix it.
Check 3: test the folk remedies people swear by
An entire folklore has grown up around beating these tools. Most of it doesn’t work, and a few of the tricks actively make the score worse. Test each one against the same 200-word base text and note the movement.
- Swapping in fancier synonyms: almost no change. Detectors score structure, not vocabulary size.
- Adding typos and casual phrasing: small, inconsistent dips. Typos register as noise, not as humanity.
- Breaking long sentences into fragments: usually the biggest single win, because it raises burstiness.
- Stacking em dashes and transition words: frequently backfires. Those are high-frequency markers in generated prose.
- Adding one specific, dated anecdote: the strongest move on the list. Names, numbers, and places pull a paragraph out of the generic register quickly.
Check 4: separate structure from prose
Many detectors score in chunks and handle headings, lists, and short paragraphs badly. Identical content written as four flowing paragraphs and as eight bullets can produce wildly different numbers. Try both with the same wording underneath. If the bulleted version scores 25 points higher, you’ve found a formatting artifact rather than a writing problem.
Check 5: run the same file twice, a day apart
Upload the identical document 24 hours later. Some tools ship model updates weekly, and some aren’t deterministic even on the same model. A swing of more than 15 points on unchanged text tells you the score carries far less certainty than the interface implies. Knowing which detectors hold steady and which drift is more useful than knowing which one advertises the highest accuracy.
What to do when a flag survives all five checks
If a detector flags your original work consistently, and you’ve confirmed it doesn’t flag pre-2022 writing or a fully rewritten draft, you finally have something concrete to present. Keep a version history. In Google Docs, the edit timeline with timestamps is the single strongest piece of evidence you can hand over. Add interview notes, source documents, and early outlines alongside it.
Screenshot the detector’s failure on your oldest sample and place it next to your flagged piece. Two images do more work than a paragraph of protest. If ownership matters over the long term, there are technical routes too, such as watermarking your text so you can prove later where a passage came from.
Where detectors belong in an actual workflow
Treat them as a smoke alarm, not a fire investigation. Run one on your own draft before you submit anything, so a client’s score never catches you cold. When a reading comes back high, use it as a prompt to reread your paragraphs: are they all the same length? Is every sentence doing the same grammatical job? Those are genuine weaknesses worth fixing regardless of what any classifier believes.
Then get back to the writing. If you want the operational version of this, one that covers picking a tool, setting a threshold, and logging results across a team, there’s a practical step-by-step detector workflow worth following. Just keep the numbers in proportion. A score is one data point from a system that is guessing, and the editor, teacher, or client arguing with you deserves to know what that system can and cannot see.

