Your inbox pings. A freelance writer has sent over a 700-word blog post. The topic is fine, the deadline is tight, but something feels off. The sentences are a little too smooth. The transitions are a little too predictable. You paste a chunk into an AI text detector. The result: 87% likely AI-generated.
Now what? Do you reject the piece? Ask for a rewrite? Or do you dig deeper?
Most guides stop at the score. That’s a mistake. AI text detectors are probabilistic tools, not lie detectors. This step-by-step workflow will show you how to use one responsibly, with concrete examples at each stage.
Step 1: Know what the detector is actually measuring
An AI text detector doesn’t read for meaning. It looks at statistical patterns: perplexity (how surprising word choices are), burstiness (variation in sentence length and structure), and token frequency. Human writing tends to be messy and unpredictable. AI writing tends to be consistent and smooth.
That’s why a single number is never proof. As we’ve explained in our guide to using imperfect AI text detectors, even the best tools produce false positives and false negatives. Treat the score as one data point, not a verdict.
Step 2: Prepare a clean, representative sample
Don’t paste the entire document. Detectors work best on continuous prose of 300 to 500 words. Longer samples can dilute the signal, and shorter samples are too noisy.
What to include and what to cut
- Remove headings, bullet points, quotes, and citations.
- Skip any text you know is human-written, like a client brief or an interview quote.
- Take a sample from the middle of the document, not just the intro. Intros are often edited more heavily.
- If the document is short (under 300 words), wait until you have more text to test.
Example: For a 1,200-word article, copy the first two paragraphs of the introduction and one full body section. That gives you a clean 400-word sample.
Step 3: Run the sample through at least two detectors
No single tool is reliable on its own. GPTZero, Originality.ai, Copyleaks, Sapling, and Writer.com all use different models and training data. They will disagree.
In one test, a paragraph written by ChatGPT-4 scored 98% AI on GPTZero but only 55% on Originality.ai. Another paragraph, written by a human journalist, scored 72% AI on Copyleaks but 12% on GPTZero. The disagreement is the point. If two detectors both flag the same text as highly likely AI, that’s a stronger signal than one outlier.
For a deeper comparison of what each tool does well, see our breakdown of ChatGPT detectors and what to trust.
Step 4: Interpret scores with context
A score of 50% does not mean “half AI.” It means the detector is uncertain. You need to understand the common failure modes.
False positives happen
Formulaic human writing often trips detectors. A legal disclaimer (“The party of the first part shall indemnify…”) might score 90% AI because it’s repetitive and low in burstiness. Non-native English speakers sometimes get flagged because their sentence structures are simpler and more consistent than native speakers’. Technical writing, like a software manual, can also score high.
False negatives happen too
AI text that has been lightly edited by a human can slip through. Ask ChatGPT to write a product description, change a few adjectives, and the detector might drop to 20% AI. That doesn’t mean the text is original.
This is why you should never rely on one score from one sample. Look for a pattern across multiple detectors and multiple samples. If the results are all over the place, you need more evidence.
Step 5: Do a manual style check
Detectors are useful, but your eyes are still the best tool. Read the sample out loud. Listen for rhythm. AI text often has a flat, even cadence. Here are specific red flags:
- Overuse of transition phrases like “Moreover,” “Furthermore,” or “In addition.”
- Vague claims without specific details: “Many experts agree” instead of “According to a 2023 Stanford study.”
- Repetitive sentence openings. Three sentences in a row starting with “The” or “This” is suspicious.
- Buzzwords that cluster together: “delve,” “tapestry,” “landscape,” “testament,” “navigate.”
- No personal anecdotes, specific dates, or odd details that a human would naturally include.
But be careful. Plenty of human writers use these patterns, especially under deadline. One red flag is not proof. Three or four in a single paragraph, combined with a high detector score, is worth investigating.
Step 6: Test a second sample from a different section
AI writing rarely appears uniformly across a long document. Maybe the introduction was AI-generated, but the body was written by a human. Or the conclusion was copied from a chatbot, but the rest is original.
Take another 300-word sample from a different part of the document. If the first sample scored 90% AI and the second scores 30%, you likely have a mixed document. That changes your next step. You might ask the writer to revise the flagged section rather than reject the whole piece.
Step 7: Decide what to do with the results
No detector gives you certainty. So your response should be proportional.
If two detectors flag the same sample as over 80% AI, and your manual check finds multiple red flags, you have a reasonable case to ask the writer about their process. Don’t accuse. Ask a simple question: “Can you tell me how you wrote this section? I’m seeing some patterns that look like AI generation.”
If the results are mixed, or the score is below 60%, treat it as inconclusive. Ask for revisions that add specific details, examples, or personal voice. That’s often more productive than a debate about the score.
For a framework on handling these gray areas, see AI writing detectors aren’t perfect — here’s how to use them anyway.
Step 8: Use other detection methods when AI text isn’t the only concern
An AI text detector won’t help you catch copied text. If you suspect someone has plagiarized a human-written article, you need a plagiarism checker or a watermarking approach. For example, you can embed invisible patterns in your own writing and later prove ownership. Our tutorial on text watermarking in Python walks through how to do that.
Also keep in mind that AI detectors are trained on older models. As newer models like GPT-4o and Claude 3 become more human-like, detection accuracy drops. A tool that worked well in 2023 might be useless in 2025. So your workflow needs to be flexible, not dependent on any single tool.
Putting the workflow into practice
Here’s the short version. Take a 400-word sample. Run it through two detectors. Note the scores. Do a manual style check for specific red flags. Test a second sample from a different section. Then decide on a proportional response based on the total evidence.
That process takes about 15 minutes. It won’t give you a binary answer, because no AI text detector can. But it will give you something more useful: a defensible, repeatable way to make a judgment call. And that’s what you actually need when you’re sitting in front of a suspicious document and wondering what to do next.

