Methodology

AI Detector Accuracy: How Our Checker Really Works

“Are AI checkers accurate?” deserves a straight answer, not a marketing percentage. This page explains exactly how our detector scores text, where it is strong, where every detector fails, and how to use the results responsibly.

Comparing tools across the industry instead? Read our full guide: Are AI Detectors Accurate?

How the scoring actually works

A scan is not one model returning one number. Several independent layers each look at the text, and the layers are deliberately capped so no single signal can dominate:

1. An AI-model classifier

A language model (run through OpenRouter) reads the text and estimates an overall AI probability, weighing signals like predictability of word choice, uniformity of sentence rhythm, hedging language, and generic phrasing. It is explicitly instructed to avoid false-positive traps: formal writing, technical domains, and non-native English patterns are not treated as AI indicators on their own.

2. Sentence-level scoring

The text is split into sentences and each one gets its own 0–100 probability, so the result highlights which specific lines look AI-written instead of issuing a single document-level verdict. For long documents, the first 30 sentences are scored individually.

3. Statistical pattern analysis

Independent of any AI model, the text is measured for burstiness (variation in sentence length — human writing varies a lot, AI clusters tightly), repeated n-grams, em-dash density, AI-typical vocabulary and phrase patterns, and structural tells like rule-of-three lists. These pattern scores are capped so they complement the classifier rather than replace it.

4. Distribution-level signals

Drawing on distribution fine-tuning research (Rosmine, 2026), which documents how LLM output diverges statistically from human writing, the detector checks distribution-level fingerprints: several consecutive sentences starting with the same word, over-reliance on starters like “The,” tokens that appear far more often in AI output than in human text, default AI-invented names, and high passive-voice density. Each finding is named in your results, with its severity.

5. Radar metrics and confidence band

The result includes a six-axis profile — AI score, burstiness, repetition, vocabulary, structure, punctuation — plus a written explanation and a confidence band. Scores of 20 or below and 80 or above are High confidence; the middle range is labeled “Low — recommend human review,” because that is the truth about mid-range scores.

What “accuracy” actually means here

A single accuracy percentage hides the two ways a detector can be wrong, and they have very different costs:

  • False positives — human writing flagged as AI. This is the damaging error in education and hiring: a student or writer gets accused of something they did not do. Polished formal writing and non-native English are the most common victims.
  • False negatives — AI writing that scores as human. Paraphrasing, humanizer tools, and heavy editing all push AI text below detection thresholds. Every detector misses some of it.

Any tool can look “99% accurate” by testing on raw, unedited AI output and clean human prose. Real-world text — edited, mixed, translated, or deliberately humanized — is where accuracy claims fall apart. That is why we publish confidence bands and named signals instead of a headline benchmark number: you can see why a score is what it is, and how much to trust it.

Why detectors disagree with each other

Run the same essay through five detectors and you will often get five different numbers. Each tool uses different models, different training data, different statistical features, and different thresholds for “AI.” None of them are measuring an agreed-upon quantity — “62% AI” has no shared definition across tools. When detectors disagree on your text, the correct conclusion is usually that the text sits in the gray zone: partially edited, mixed authorship, or simply short enough that the statistics are thin. That calls for human judgment, not a sixth detector.

Responsible use

  • Never treat a score as sole evidence in an academic-integrity or hiring decision.
  • Give more weight to results on longer texts — very short passages do not contain enough signal for any detector.
  • Apply extra caution to flagged writing from non-native English speakers, who are disproportionately affected across the industry.
  • Use the sentence-level highlights to have a concrete conversation, and pair scores with process evidence like drafts and version history.
  • Respect the confidence band: a mid-range score labeled “recommend human review” means exactly that.

See the methodology on your own text

Paste up to 500 words free — no account — and get the full result: per-sentence probabilities, named signals, radar metrics, and an honest confidence band.

Try the free AI checker

Frequently Asked Questions