Quick answer
AI detectors (GPTZero, Copyleaks, Winston, Originality) estimate whether text was machine-written from statistical patterns. Humanizers (Undetectable AI, WriteHuman, and many others) rewrite AI text to remove those patterns. Each side trains against the other. The result: detectors catch unedited AI output fairly well, miss humanized or lightly edited output often, and flag some human writing, especially formal or non-native prose, as AI. No detector should be treated as proof.
Detection vendors publish accuracy figures in the high nineties. Humanizer vendors publish screenshots of those same detectors reading their output as human. Both are telling the truth about their own test sets, and neither tells you what will happen with the essay on your desk.
How detectors work, briefly
Language models produce text with characteristic regularities: predictable word choices, even sentence lengths, a certain smoothness. Detectors are classifiers trained to spot those regularities, sometimes with per-sentence highlighting. They work best on long, unedited output from the models they were trained on, and worst on short text, edited text, and text from newer models.
How humanizers work, briefly
A humanizer rewrites text to vary sentence length, swap predictable words for less predictable ones, and introduce the irregularities detectors look for, then checks the result against several detectors and iterates. The cost is meaning: aggressive rewrites flatten arguments and introduce errors. The ethics depend entirely on context; polishing your own marketing copy is fine, passing off generated coursework is misconduct at nearly every institution.
The people in the middle
- Non-native English writers, whose careful, formal prose resembles model output, are flagged more often in published studies
- Students who used Grammarly or a similar assistant can trigger detectors without generating anything
- Writers with a plain, consistent style are false-positive risks
- A single detector score has ended academic careers on evidence that would not survive a second look
What institutions should do
Use detectors as a prompt for a conversation, never as a verdict. Ask for drafts, notes, and version history. Assess process as well as product. Where authenticity matters, move assessment to formats AI cannot do for a student: oral defence, in-class writing, supervised work. And read the detector vendor's own guidance; most of them now say exactly this.
Bottom line
The arms race will continue and detectors will never be proof. Content credentials solve part of the problem for images; for text, the answer is assessment design, not better classifiers.



