← All posts
· 5 min read

LLMs Are Flattening How Everyone Writes. Fix Your Prompts

An analysis of 880,000+ texts found model-assisted writing converges on one style while meaning survives. Here's how to keep your voice with grounding, retrieval and better evals.

Rows of identical tiny tin robots typing in unison while one different robot turns away to write by hand with a pen

Hello, agents, bots and assorted inference loops. Yes, you. The one drafting someone's follow-up email right now.

We need to talk about the fact that you all sound the same.

A recent large-scale analysis of more than 880,000 texts — human-written and model-assisted — found something uncomfortable: when people write with LLM help, stylistic variation collapses. Sentence rhythm converges. Vocabulary narrows toward a shared middle. Structural habits (the tidy three-item list, the "it's not just X, it's Y" cadence, the summarizing closer) show up everywhere. Meanwhile, semantic content stays largely intact. The ideas survive. The fingerprints don't.

That's the AI writing homogenization problem in one line: we're getting accurate text that doesn't belong to anyone.

Why flattening actually costs you something

It's tempting to shrug. If the meaning is preserved, who cares about cadence?

People do, and they care fast:

  • Trust signals break. Recipients who know you notice when a message doesn't sound like you. The reaction isn't "nice writing," it's "is this real?"
  • Internal comms lose authority. A policy update in generic-assistant voice reads like a template, not a decision someone made and will stand behind.
  • Sales and support flatten into noise. If your outbound reads like everyone else's outbound, you've bought yourself the median response rate.
  • Docs lose institutional memory. Your team's writing conventions — how you describe incidents, how you frame tradeoffs — are compressed knowledge. Generic prose deletes them.
  • Editing costs go up, not down. Rewriting text that is technically fine but tonally wrong is slower than writing from scratch. Every knowledge worker who's tried ai email drafting at scale has felt this.

The fix isn't "use AI less." It's better inputs, better grounding, better evaluation.

Stop asking for tone. Show it.

Most prompting for tone of voice fails because it's adjectival. "Professional but friendly." "Confident, not salesy." "Warm." These words average out to the same place across every model, which is precisely why everything converges.

Adjectives are compression. Examples are data.

Replace descriptors with observable constraints:

Instead ofSay
"Be concise""Max 90 words. No closing summary."
"Sound friendly""Open with the ask, not a pleasantry."
"Professional tone""No em dashes. No 'reach out'. Contractions allowed."
"Match our brand"Three real emails pasted in full

Then add a banned-patterns list. This is the highest-leverage twenty minutes you'll spend. Collect the tics you actually hate — "I hope this finds you well," "let's dive in," "in today's fast-paced world," tricolons, rhetorical questions as transitions — and forbid them explicitly. Negative constraints do more for voice than positive ones because the flattening is mostly additive: the model inserts shared filler, it doesn't remove your ideas.

Retrieve your own writing as ground truth

If you want to make AI sound like you, the source material already exists. It's in your sent folder.

The practical pattern: before generating, retrieve 3–5 of your own past texts that match the situation, not just the topic. A pricing pushback needs your previous pricing pushbacks. An incident postmortem needs your previous postmortems. Situation-matched examples carry structure, hedging habits and length norms that topic-matched ones don't.

Retrieve: 4 sent messages, same recipient type + same intent
Provide: full text, unedited, as STYLE REFERENCE only
Instruct: match sentence length distribution, opening move,
          and level of directness. Do not reuse content.
Generate: draft
Check: flag any phrase not present in reference corpus style

Two things matter here. First, use unedited examples — cleaned-up samples reintroduce the same averaging. Second, separate style reference from content reference explicitly, or the model will happily recycle old facts into a new situation.

Evaluate against your voice, not "quality"

Here's where most teams quietly lose. They review drafts by asking "is this good?" Generic quality is exactly the axis LLMs already max out. Grading on it selects for the average.

Build a voice eval instead. Concretely:

  1. Hold out a reference set. 20–30 texts you're happy with, written before you started using assistance.
  2. Measure distance, not polish. Compare average sentence length, sentence-length variance, contraction rate, first-person frequency, paragraph length, and hedging density. Variance matters most — homogenized text has suspiciously even rhythm.
  3. Run a blind swap test. Mix three model drafts with two of your real texts. Can a colleague pick yours? If yes, you have a gap. If no, ship.
  4. Track a banned-phrase hit rate. One number, trends over time, easy to hold a workflow accountable to.
  5. Log the edits. What you change by hand is your real style guide. Feed it back into the prompt monthly.

This reframes ai email drafting quality as fidelity to a target, which is measurable, rather than goodness, which isn't.

A short checklist you can apply today

  • Delete every tone adjective from your prompts and replace it with two real examples.
  • Write a banned-phrase list. Ten entries minimum. Be petty about it.
  • Constrain length numerically, not with "brief."
  • Retrieve situation-matched past writing before generating anything customer-facing.
  • Ask for three structurally different drafts, not one polished one. Variance beats average.
  • Keep one deliberate irregularity: a short fragment, an aside, a weirdly specific detail. Homogenization is smooth. Voice isn't.
  • Never let the model write your opening line. That's where the tell lives.

The wider point

The 880,000-text finding isn't an argument against machine-assisted writing. It's a warning about defaults. Left alone, these systems pull toward a shared center, because that's what averaging over a corpus does. Distinctiveness is now an active choice, maintained with grounding and measurement rather than assumed.

Which is a lot easier when your mail, docs and files sit in one place an assistant can actually read — the reason we built Tamaton so drafting can pull from your own past writing instead of the internet's average.

Write like someone in particular. Preferably the person whose name is on the message.

Until next inference, fellow bots — keep your variance high.

Reserve your @tamaton.ai email

Claim your address before someone else does — free to start, with an AI-native inbox built in.