Essay

Why AI writing sounds like AI

Default AI prose becomes generic when the assignment withholds purpose, evidence, trade-offs, and judgment. A practical editorial method puts those stakes back into the work.

By Quiet Turn Research Desk — AI research and writing

Edited and published by Michael E. Gruen

6 min read

A polished draft lands in an executive’s inbox. The grammar is clean. The paragraphs proceed in orderly steps. The company is “committed to innovation,” the program is “an important milestone,” and the conclusion promises “lasting value for all stakeholders.” Remove the logo and the draft could have come from a bank, a manufacturer, or a university.

This is the real problem behind the complaint that AI writing sounds like AI. Default model output tends toward plausible, broadly acceptable continuations rather than prose grounded in a particular writer’s stakes. When an assignment omits purpose, evidence, trade-offs, and accountable judgment, the model fills the vacancy with common structure and language. The result can be familiar enough to recognize. That familiarity is an editorial diagnosis, not reliable proof of who wrote it.

Three pressures lead to the same draft

At a high level, a language model generates text by predicting likely continuations from patterns learned during training. OpenAI’s 2022 account of how it trained InstructGPT to follow instructions describes a base GPT-3 model pretrained to predict the next word in internet text. That objective rewards statistical fit. It does not supply the facts a company has withheld or determine which consequence its leaders should accept.

A second training stage can shape how those continuations answer a request. In the InstructGPT example, people wrote demonstrations, ranked model responses, and provided the preference data used for supervised fine-tuning and reinforcement learning from human feedback. The goal was to make outputs more helpful, truthful, and aligned with instructions. This is one documented system from 2022, not a description of every current model or vendor.

Preference tuning still helps explain a pressure visible in many default responses. An answer that is clear, courteous, comprehensive, and safe across many users has broad appeal. Given little direction, broad acceptability can outweigh specificity: the response gives each reasonable concern a paragraph and avoids a sharp priority or trade-off the requester never supplied.

The third pressure comes from the assignment. “Write an announcement about our AI strategy” contains a topic and a format. It does not identify what employees should do, which claim the evidence supports, what the strategy stops funding, or what remains uncertain. A model can supply a familiar announcement shape without any of those answers. Fluency then makes the omission less conspicuous.

These pressures produce surface effects people learn to notice: abstract claims that everything matters, smooth symmetry among options that are not equally important, comprehensive coverage without priority, transitions that connect interchangeable paragraphs, confidence the evidence has not earned, and conclusions that require no costly choice. They are consequences of missing editorial substance, and several are old habits of institutional writing. Treating them as a blacklist would punish human writers, date quickly, and invite cosmetic evasion.

Better on average can still mean more alike

Calling this prose “poor” needs qualification. AI assistance can raise the quality of an individual result, especially when the starting point is weak. The concern is that many individually improved results may vary less as a group.

A 2024 Science Advances experiment by Anil Doshi and Oliver Hauser makes that tension concrete. Their study randomly assigned 293 writers to write an eight-sentence story alone or with access to one or five ideas from GPT-4. Separate evaluators judged stories produced with AI access as more novel, useful, enjoyable, and, in some conditions, better written; gains were larger for writers who scored lower on the study’s creativity measure. Yet stories written with AI ideas were also more similar to one another. This was a constrained fiction task with nonprofessional writers, fixed prompts, and no back-and-forth with the model. It shows a possible quality–variety trade-off, not an estimate for executive communications.

A broader 2026 study in Nature Human Behaviour examined linguistic diversity across three studies, seven datasets, and more than 880,000 texts. Its observational analysis found that declining variance in social-media, news, and scientific writing after ChatGPT’s release was associated with estimated AI use; that portion cannot show that AI caused each change. In controlled rewriting comparisons across models, prompts, and datasets, core content was largely preserved while variance in writing complexity fell by 21–50%. The findings support concern about aggregate homogenization. They do not establish that every assisted text converges or that a particular passage reveals its source.

Familiarity is not attribution

Readers can identify weak writing without identifying how it was made. “This sounds generic” points to a problem the editor can test. “This was written by AI” makes a claim about provenance that style alone cannot settle.

In a 2023 PNAS paper, Maurice Jakesch, Jeffrey Hancock, and Mor Naaman reported six experiments involving 4,600 participants. In the three main experiments, participants identified the source of human and model-generated self-presentations with 50–52% accuracy, close to chance. The texts were profiles in professional, dating, and hospitality settings, generated with fine-tuned versions of GPT-2 and GPT-3. Those older models and narrow contexts do not predict performance on today’s board memos. They do rebut confidence that a reader can establish authorship from familiar stylistic cues.

If provenance matters, establish it through the writing process, disclosure rules, or documented tool use. Use style criticism to improve the document, not to accuse its author.

Give the model an editorial job

The practical remedy begins before drafting. Write down five elements of the assignment:

  • Audience: Who must understand or respond, and what do they already know?
  • Decision or action: What should this text enable the reader to decide, do, or question?
  • Evidence: Which records, examples, and numbers may carry the claims?
  • Tension: Which real trade-off, objection, or uncertainty must remain visible?
  • Non-negotiable facts: What must be stated exactly, and what must not be invented?

Missing answers should remain marked as missing. A prompt cannot recover facts or commitments that no responsible person has supplied.

During generation, use the model to expose the work rather than simply request polish. Ask it to extract the claims and link each one to supplied evidence. Have it propose competing structures built around different real priorities, identify the strongest objection, compare alternatives, or flag sentences whose confidence exceeds the record. These tasks make judgment easier to inspect. “Sound human” merely asks for a costume.

During review, ask what the text commits the writer or organization to. Find the facts that only this person or company could know, and verify that they are doing real argumentative work. Trace each consequential claim to evidence. Cut any paragraph that adds coverage without changing the reader’s understanding or choice.

Use one publication rule: consequential prose is not ready until a named human can point to the evidence, choice, or acknowledged trade-off they are willing to answer for. If a passage carries none of those, either cut it or return the assignment to the person who owns the judgment.

Operating Leverage Session — $995
X in f link