The writer's problem
A poet posting an unfamiliar short poem online, or a reading group swapping stanzas without bylines, is testing something readers now do constantly without a rubric: guessing whether a poem was written by a human. A paper in Scientific Reports by Brian Porter and Edouard Machery put that guess to a controlled test, using short poems generated 'in the style of' ten well-known poets, including Shakespeare, Dickinson and Chaucer, set alongside those poets' own published work.
What the documents show
The paper's own methods report two Prolific experiments. In the first, 1,634 participants each read ten poems attributed to one poet: five genuine and five produced by ChatGPT running on GPT-3.5, prompted to write in that poet's style. Accuracy at guessing authorship was 46.6 percent, below the 50 percent expected by chance, and participants were more likely to call an AI-generated poem human-written than to call a real poem human-written. In a second experiment, 696 participants rated ten poems on qualities including rhythm and beauty; the AI-generated poems scored higher, which the authors say helped drive their misidentification as human work. A second document complicates a citation the newer paper makes in passing. It states that people 'could reliably distinguish the poetry of GPT-2 from human-written poetry,' pointing to an earlier study by Nils Köbis and Luca Mossink. That paper's own results were mixed: detection failed only when a curator had already picked the most convincing AI sample, and succeeded when a poem was chosen at random. A one-line citation compressed a two-condition finding into a single direction.
The editorial choice
An editor should not read the newer paper as proof that disclosure no longer matters because readers cannot tell the difference. The study measures perceived humanness among non-expert Prolific participants judging short poems stripped of context, not literary judgment by an informed reader or an editor familiar with a poet's body of work. Editorially, the safer response is to keep authorship information attached regardless of whether a casual reader could detect a poem's origin unaided, since the paper's own design shows accuracy shifts with presentation and comparison set rather than tracking a fixed human tell.
What stays with the author
Choosing which poem to draft, revise, or publish under one's own name is not a task either study measures or could replace. The favorable ratings the Scientific Reports paper documents for rhythm and beauty describe an aggregate heuristic among readers unfamiliar with the source, not a verdict on a poem's standing within a poet's own project or a wider tradition. Neither paper resolves what happens when a reader is told a poem's origin in advance rather than left to guess it blind.
- Would these readers still prefer the same poems if the authorship were disclosed upfront?
- Does sounding human measure anything a working poet or editor should optimize for?
- What does a ten-poet comparison set miss about styles less familiar to general readers?
Both studies are lab measurements of snap judgments on short, decontextualized text: useful for showing how fragile a 'human tell' can be, not a substitute for judging what a poem is worth.
Follow the source.
The paper's own methods and results report two Prolific experiments (N=1,634 then N=696) in which non-expert readers identified AI-generated poems at below-chance accuracy (46.6 percent) and rated them more favorably on rhythm and beauty.
Source date: 14 Nov 2024 · Retrieved: 16 Sept 2026
This earlier GPT-2 study's own results show detection succeeded only when an AI poem was chosen at random, not when a human curator had picked the most convincing AI sample, a more mixed finding than the newer paper's citation of it suggests.
Source date: Not established · Retrieved: 16 Sept 2026
Site publication is not established by an event date. Original record ID: 0030-bf-091. This local design review does not change its editorial status.