LLMs Are Quietly Making Everyone Write the Same
The takeaway first: using an AI assistant to write does not just polish your text — it measurably makes your text look like everyone else's, and the effect is already visible at population scale. A team at the University of Southern California and collaborators, publishing this week in Nature Human Behaviour, analyzed more than 880,000 texts written between 2018 and 2024 and found that the spread of LLM-based writing help — ChatGPT and Claude now serve more than 800 million users — correlates with a steady collapse in stylistic diversity across arXiv abstracts, news articles, and Reddit stories alike. In a controlled test, rewriting human texts with GPT-3.5, Llama 3 70B, or Gemini Pro compressed the variance in writing complexity even when the meaning stayed intact. The same paper shows the signals that let readers infer a writer's age, personality, or values are fading — and skewing toward one "default persona". Because rewritten text then re-enters training corpora, the effect compounds. This is not a style debate; it is a measurable structural change in human language.
Study 1: Style Variance Collapses Across 880,000 Texts
Study 1 asked whether LLM adoption correlates with homogenized writing, in two steps. First, observational data: three longitudinal corpora — about 80,000 arXiv abstracts, 379,600 Patch News articles, and 318,000 Reddit creative stories, spanning 2018 to 2024. The team used a detection tool called Binoculars to flag AI-written text, then tracked the variance of a "writing complexity" metric over time.
Results were consistent. After ChatGPT launched, variance dropped significantly and persistently in arXiv and Reddit. Adoption rate predicted later variance changes — Granger causality held for arXiv; Reddit showed the same pattern with a lag of about 20 months, which the authors attribute to slower tool adoption among a broader user base. Patch News also dipped sharply at launch, but the causal signal was weaker; professional editing and institutional norms may buffer the direct effect, though subtle influence can still leak in through changing norms.
Second, a controlled experiment: 1,000 pre-ChatGPT human texts drawn from Reddit and arXiv were rewritten by three LLMs (GPT-3.5, Llama 3 70B, Gemini Pro) with neutral prompts like "improve the grammar". 87% of the rewrites kept over 0.95 similarity to the original meaning — yet the variance of writing complexity shrank significantly. Meaning preserved, style flattened.
Study 2: Identity Becomes Harder to Read — and Skews
Study 2 asked a sharper question: can you still tell who wrote a text? The team assembled text corpora paired with psychometric data — age, gender, Big Five personality, empathy, and moral values — trained classifiers to predict these traits from text, and compared prediction accuracy on original versus LLM-rewritten versions.
Two findings. First, accuracy dropped by about 6% in F1 on average across the six trait groups, with age hit hardest (F1 from 0.351 to 0.260). The drop is real but not total: classifiers still beat random chance — AI weakens identity signals, it does not erase them.
Second, the change is not random noise; it points in a consistent direction. Rewritten text was systematically judged as older, more moral, less empathetic, less extroverted, and more open and agreeable — across all three models and across different prompt phrasings. In other words, LLMs are not merely scrubbing identity; they are injecting a bias toward a shared "default persona". Individual authors are being blended toward the same statistical average.
Study 3: Classic Word–Person Links Are Being Selectively Erased
Psycholinguistics has accumulated well-known associations: extroverts use more positive-emotion words, openness correlates with complex vocabulary, and so on. Study 3 tested whether these links survive LLM rewriting. In original texts, the team reproduced many classic findings. After rewriting, several key associations disappeared: the link between extroversion and pronoun use, between loyalty and friend-related vocabulary, and between age and future-oriented words. Others — such as neuroticism and negative-emotion words — survived.
That selectivity is worse than uniform noise. Every discipline that reads people through word choice — clinical screening, social science, marketing personas — now has to wonder which associations still hold and which are quietly dead. Researchers can no longer assume their lexical markers measure what they used to measure.
Why This Compounds: The Training-Data Feedback Loop
LLM-polished text does not stay in private documents. It flows into online corpora — reviews, forums, papers, news — and those corpora train the next generation of models. The homogenizing bias is therefore self-reinforcing: each model generation reads a corpus fuller of AI-flavored prose and pushes the average further. Public open-training runs like Marin 535B make this pipeline visible: the taste of the training data becomes the taste of the model.
The structural stakes are concrete: applied fields that depend on lexical diversity — detecting depression markers from text, personalized messaging, hiring screens that favor "polished" applications, and cultural preservation for languages and dialects whose markers are being flattened — all receive progressively noisier signals. And because language and thought are entangled, a homogenized language may quietly constrain cognitive flexibility and creativity — the Orwell scenario from 1984, arriving not by decree but by "can you just help me polish this". Why this convergence is almost inevitable, and where the concrete costs land, is the focus of a companion analysis.
What to Do: Keep Human Text in the Loop
None of this argues against LLM writing tools — clarity and access are real wins. It argues for design and workflow choices:
- For researchers: revalidate lexical markers before trusting them; "this association survived an LLM rewrite" should become a prerequisite for any new corpus-based finding.
- For teams shipping AI writing features: verifiable watermarking is becoming standard for AI content — keep and label the human-authored original, because the unpolished draft is the only training signal that still carries human diversity. Anthropic's funding of open benchmarks for AI wellbeing is the same instinct applied to a different impact channel.
- For platforms and publishers: keep authentic human corpora out of the homogenization loop — the less AI-flavored text in the next training set, the slower the drift.
- For writers: treat "polish" as a lossy operation. Keep the unpolished draft, review edits deliberately, and question edits that nudge your voice toward the default.
FAQ
Q: How do we know AI writing is actually reducing language diversity?
A: Observational data across 880,000+ texts (arXiv, Patch News, Reddit, 2018–2024) show writing-complexity variance fell after ChatGPT launched, with adoption predicting the drop. A controlled experiment rewrote 1,000 human texts with three LLMs and reproduced the variance compression while 87% of rewrites kept over 0.95 semantic similarity.
Q: Does AI writing erase identity signals completely?
A: No. Trait-prediction accuracy dropped about 6% in F1 (age: 0.351 to 0.260) but stayed above chance. The bigger problem is systematic bias — rewritten texts skew toward an older, more moral, less empathetic "default persona", consistently across models and prompts.
Q: What is the actual long-term risk?
A: A feedback loop: AI-flavored text enters online corpora, trains the next models, and pushes the average further. Fields that depend on lexical diversity — depression screening, hiring, marketing, cultural preservation — get progressively noisier signals.