AI detectors analyze patterns that correlate with generated language. They do not observe authorship directly. The result is an estimate influenced by the model, threshold, text length, language, and writing style.
What the score means
A percentage is best understood as the detector’s confidence under its own assumptions—not the percentage of words written by AI. Different systems can score the same passage differently because they use different data and thresholds.
Why false positives happen
Formulaic, highly edited, concise, translated, or non-native-English writing may resemble statistical patterns found in generated text. Short samples provide less evidence and often produce less stable results.
A responsible review process
- Use a sufficiently long, representative sample.
- Compare multiple passages rather than one paragraph.
- Look for unsupported claims and invented citations.
- Ask the author to explain reasoning and sources.
- Treat detector output as one input, never proof.
Focus on quality and provenance
The most useful question is often not “Was AI involved?” but “Is this work accurate, original, transparent, and accountable?” That standard holds whether a writer used a model, an editor, or neither.