Work Out How Often a Flagged Essay Was Written by a Student
Three numbers together answer the only question that matters when a flag appears: how often is a flag right? Take a hundred pieces of work. Suppose ten of them really are machine-written, the tool catches every one of them, and it wrongly flags one percent of the ninety human pieces. That gives ten correct flags and about one wrong flag: roughly one flag in eleven lands on a student who wrote their own work, and that is with a tool that never misses, which no evaluated tool achieves. Now change the assumptions to realistic ones. If only five in a hundred are machine-written and the false positive rate is the nine percent OpenAI measured on its own classifier, you get five correct flags and about nine wrong ones, so most flags are mistakes. Counting in whole documents rather than percentages, the way natural-frequency research shows people reason best, makes this visible in a minute on paper. You can now take a false positive rate, a sensitivity figure and your own estimate of how much AI writing is in a stack, and state what share of the flags you should expect to be wrong before you act on any of them.
Sensitivity and Specificity in Diagnostic Testing
Sensitivity and specificity are statistical measures used to evaluate diagnostic test accuracy: sensitivity is the proportion of diseased individuals who correctly test positive, while specificity is…