Read the evidence / Source-led record

AI-text detectors flagged most non-native writers' essays

A Stanford study found seven GPT detectors misclassified 61 percent of non-native English essays as AI-written, versus near-perfect accuracy on native essays.

The writer's problem

A non-native English writer submitting an essay, a query letter or a manuscript sample to an editor or a program that screens submissions with an AI-text detector faces a specific risk: being flagged as a machine even when they wrote every word. A Stanford study by Weixin Liang, Mert Yuksekgonul, Yining Mao, Eric Wu and James Zou measured that risk directly rather than assuming it.

What the documents show

The paper's own results state that seven widely used GPT detectors were tested on 91 TOEFL essays written by non-native English speakers, sourced from a Chinese educational forum, alongside 88 US eighth-grade essays from the Hewlett Foundation's public dataset. The detectors showed near-perfect accuracy on the native-authored essays but misclassified the TOEFL essays as AI-generated at an average false-positive rate of 61.22 percent; all seven detectors unanimously flagged 18 of the 91 TOEFL essays. Prompting ChatGPT to enrich the TOEFL essays' word choice cut the false-positive rate to 11.77 percent, while simplifying the native essays' word choice raised their false-positive rate from 5.19 percent to 56.65 percent, suggesting the detectors were responding to linguistic variety rather than authorship as such. The peer-reviewed version, published in Patterns on 10 July 2023, carries the same findings under formal review, confirming this was not left as an unreviewed preprint claim.

The editorial choice

A journal, program or contest that screens submissions with a detector should not treat a flag as evidence of AI authorship on its own, particularly for writers whose English is a second language, since the paper's own design shows the same detectors respond to lower lexical variety regardless of its source. Editorially, a flagged submission from a non-native writer warrants a human read before any consequence follows, not automatic rejection.

What stays with the author

The study does not resolve whether any specific detector in current commercial use behaves the same way, since it tested named tools as they performed in 2023, and it does not tell a writer how to prove authorship beyond what a human editor's judgment can establish. Choosing whether to write in a plainer or more elaborate register, and standing behind the resulting text, remains the writer's own decision, not one a detector score settles.

  • Does a screening process rely on detector output alone, or is a human reader also checking flagged submissions?
  • Would this bias apply to the specific detector a given platform or contest actually uses today?
  • What obligation does a screener have to a writer wrongly flagged, given a documented, unequal error rate?

Together the two documents establish a specific, measured disparity in one set of tools at one point in time, reason for caution in how a flag is used, not proof that every detector or every writer is affected the same way.

Follow the source.

GPT detectors are biased against non-native English writers ↗

The paper's own results report that seven GPT detectors misclassified non-native TOEFL essays as AI-generated at an average 61.22 percent false-positive rate versus near-perfect accuracy on native-authored essays, and that word-choice prompting shifted the rate in both directions.

Source date: 6 Apr 2023 · Retrieved: 16 Sept 2026

GPT detectors are biased against non-native English writers (Patterns, peer-reviewed record) ↗

The PubMed Central record confirms peer-reviewed publication in Patterns, volume 4 issue 7, on 10 July 2023, establishing this was not left as an unreviewed preprint claim.

Source date: 10 Jul 2023 · Retrieved: 16 Sept 2026

Site publication is not established by an event date. Original record ID: 0030-bf-097. This local design review does not change its editorial status.

Keep following the question

Next on your desk.