
Test AI detectors on the writing people actually submit
A detector’s strong result on generic AI prose can weaken on style-conditioned text. Evaluate the intended genre and review both kinds of error.
A detector that recognises generic AI prose has passed one test. It has not yet passed the test a publishing workflow gives it.
Epoch AI’s comparison makes the gap concrete. Across 297 plainly prompted AI passages, each of three detectors missed at most two. On a separate set of 297 passages generated with style examples, the detectors missed between 30 and 53. The result concerns the tested versions and conditions, not an enduring ranking of the products.
I would make that second condition part of an ordinary procurement evaluation. A tool intended for edited, specialised writing should encounter edited, specialised writing before anyone approves a consequential use for its verdict.
The difficult input is part of the product
It is tempting to put a near-perfect headline score in the main presentation and difficult cases in a red-team appendix. That arrangement makes the ordinary test look representative and the realistic variation look exceptional.
A writing assistant can receive examples, house-style guidance or domain-specific instructions during legitimate use. A detector may therefore encounter prose that differs substantially from a one-line prompt’s output without anyone attempting deception. The evaluation should cover the workflow’s actual conditions, while respecting rights and permission to use the test material.
The study does not show that all detectors are useless. In its human set, Pangram and GPTZero flagged none of 495 passages, while Originality.ai flagged 19. Those observations are useful too. They are estimates on a particular corpus, and zero observed false positives is not a guarantee that future human writing will never be flagged.
Decide what a mistake would do
For a hypothetical editorial desk, I would begin with the decision the detector is supposed to support. Is it selecting submissions for further review, checking a declared process, or being proposed as grounds for rejection? Those uses impose different costs when the result is wrong.
Then build a permitted evaluation set reflecting the desk’s genres and typical document lengths. Keep known human provenance and known AI generation separate. Include the kinds of assistance that the publication actually allows. A mixed workflow should not quietly become a binary label just because the dashboard prefers two colours.
The output I want is an error table by condition, with examples a reviewer can inspect. Which human work was flagged? Which generated work was missed? What changes when the genre changes? A single overall accuracy percentage can conceal the part of the workflow where the tool is least useful.
Epoch’s own reporting shows why definitions matter: treating a detector’s “Mixed” label as detection rather than non-detection changes some reported miss rates. I would decide that interpretation before comparing suppliers, then preserve it with the result.
This is an evaluation proposal, not a reproduction of the study. Its corpus and detector versions have limits, and I have not independently rerun the services. The published work tests one conditioning method; it does not exhaust how writing or detection can change.
I would also keep the editorial policy outside the classifier. Permission to use AI, disclosure of assistance and factual reliability are decisions about process and content. A detector score cannot resolve them by itself.
Approve an AI detector only for a defined review task, after testing the genres and assistance conditions that task will actually encounter.


