A client sent me a 900-word blog draft in March and asked a blunt question: did her new freelancer use ChatGPT? I ran the draft through three detectors. One reported 97% AI-generated. The second said 14%. The third timed out. Same text, same afternoon, three answers.
That afternoon shaped the workflow below. It is a six-step process for running a chat gpt detector on real writing without fooling yourself, and it takes about twenty minutes for a normal document. It works whether you are grading essays, editing freelance copy, or screening job applications.
Step 1: Write Down the Facts You Already Have
Detectors do not exist in a vacuum, so start with context instead of a score. Before you paste anything, note three things:
- Who wrote the text, or claims to have written it
- When it was written, and whether the writer had AI access at that point
- What decision hinges on the result: a grade, a payment, a job offer
That third item changes how you read the output. A 60% AI score on a casual newsletter draft means nothing. The same score on a doctoral thesis chapter is worth an hour of investigation.
It also helps to understand how these tools reach a verdict before you lean on one. Detectors lean on signals like perplexity and burstiness, not on some secret archive of chatbot sentences. That is exactly why two tools can disagree so violently about identical text.
Step 2: Calibrate the Tool With Text You Control
Never trust a detector you have not tested. Spend five minutes feeding it two known samples: roughly 200 words you wrote yourself, and roughly 200 words you deliberately generated with ChatGPT on a topic you know well.
Here is what a reasonable calibration looks like on a decent tool:
- Your own writing: 2% to 15% AI-flagged
- Raw ChatGPT output: 85% to 99% AI-flagged
If your own writing comes back at 70%, the tool is unreliable for your purposes, or your style happens to be unusually uniform. Either way, you now know. A structured set of checks for testing an AI writing detector before you trust its verdict will catch problems a single test run misses.
Step 3: Break the Document Into Chunks
Whole-document scores hide the interesting parts. Detectors average everything into one number, and averages are terrible at telling you which paragraphs deserve a second look.
A real example
A 1,400-word university essay scored 43% AI overall. Useless on its own. Chunked into six sections, the picture changed fast: intro 91%, first argument 12%, second argument 8%, case study 66%, counterargument 15%, conclusion 88%.
The student had clearly written the analysis and used a chatbot for the bookends. That is a very different conversation from “you cheated.”
Chunk at natural boundaries, 150 to 250 words each. Headings, paragraphs and argument blocks all work. Do not split mid-sentence, because detectors need enough text to calculate anything meaningful.
Step 4: Read the Sentence Highlights, Not the Headline Percentage
Most detectors color-code individual sentences in red, yellow or green. That layer is usually more informative than the number above it, because a single suspicious sentence tells you what to look for by hand.
What you will notice after a dozen runs is that flagged sentences share certain shapes. They also appear in text that no chatbot touched. Common false positive triggers include:
- Bullet lists, where every line is short and structurally identical
- Technical definitions and legal boilerplate, which are repetitive by design
- Work by non-native English speakers, whose phrasing can look statistically flat
- Heavily edited drafts that have been smoothed into uniform rhythm
That last one trips up a lot of editors. An AI text detector will flag polished human writing with the same confidence it flags a chatbot, so treat red sentences as leads rather than proof.
Step 5: Cross-Check With a Second Tool Built Differently
Run the chunks through a second detector that uses a different method. Some tools look at token predictability, some at sentence length variance, some at phrasing patterns borrowed from known model outputs. When two unrelated approaches agree on the same paragraph, your confidence should rise.
When they split, weight the agreement more than the disagreement. In my own tests, paragraphs both tools flag over 80% turned out to be generated about nine times out of ten. Paragraphs where one says 90% and the other says 10% are almost always noise. If you want a repeatable order of operations for this, this step-by-step detector workflow walks through the same sequence with screenshots.
Step 6: Decide Based Only on What You Can Explain
Before you act, write one sentence: “I believe this section was AI-assisted because the sentences here average 11 words while the rest of the document averages 22, and the vocabulary shifts noticeably.”
If you cannot finish that sentence with something concrete, you do not have evidence. You have a number. Do not fail a student, reject a candidate or withhold payment on a number.
Three Situations You Will Actually Face
You are grading student work
Chunk by paragraph, photograph the sentence highlights, and compare the flagged sections against the student’s earlier submissions. A sudden jump in polish between week 3 and week 4 is more telling than any percentage.
You are editing a freelancer’s draft
Look for the seams. AI-assisted drafts often have a fluent opening and a vague middle where specifics should be. Ask for the source of two claims. Honest writers produce them in minutes.
You are screening applications
Skip the detector for the first pass and read for specificity instead. Cover letters that name a project, a metric and a mistake tend to be human. So do the ones with an odd, memorable detail.
What No Detector Catches Yet
Lightly edited AI text slips past almost everything. A writer who generates a draft and then rewrites 30% of it in their own voice will usually score under 20%, even though a model did most of the work. Spotting that kind of AI slop takes judgment about substance, not a percentage.
There is one exception worth knowing about, and it applies to your own writing rather than other people’s. If you embed an invisible watermark in text before you publish or share it, you can later prove where a paragraph originally came from, regardless of what happened to its style along the way. Text watermarking in Python takes about thirty lines of code to set up and settles ownership questions that detectors cannot touch.
For everything else, the honest position is this: a chat gpt detector is a flashlight, not a verdict. Point it at chunks of text, calibrate it first, read what it lights up, and keep your own eyes on the page. The tools will keep changing. The habit of looking closely will not.

