What the Average AI Baseline Hides: Where the 5 Models Actually Diverge
We went back into the corpus behind our Average AI baseline and pulled the per-model scores for its 80-sample English slice. Two axes are nearly identical across models. Expressiveness spreads 61 points. Here's the real breakdown, and what it means for your Style Profile.
By Emmanuel
The "Average AI" radar chart on your Writing DNA Snapshot flattens something worth unpacking. It shows one shape, labeled Average AI, next to yours. The implication is that AI models write roughly the same way by default.
On two of the six axes, that's true. On one of them, it's badly wrong.
We went back into the same 320-sample corpus we built for how we measure average AI and pulled the English-language scores apart by model instead of averaging them together. The result sharpens something we suspected but hadn't quantified: the "median AI" isn't a single personality. It's five different personalities that happen to agree on two things and disagree sharply on a third.
TL;DR
- Five AI models (Claude Opus 4.6, Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.2, Gemini 3 Pro), measured on the same six stylometric axes, scored on the80-sample English sliceof our 320-sample corpus: vocabulary range and consistency cluster within 13-16 points across all five models. Sentence complexity, formality, and conciseness show moderate 20-29-point spreads. Expressiveness blows the others out at a 61-point spread, from Gemini 3 Pro (39) to Claude Opus 4.6 (100, the scale ceiling). The same expressiveness gap shows up in French (56) and Spanish (60), so it isn't an English artifact. If your Style Profile is calibrated to "Average AI," the override on expressiveness needs to be the strongest one you apply, because that's the axis where "the AI default" varies most by which AI you're actually using.
The Setup: Same Corpus, Different Cut
Full methodology is in how we measure average AI: 320 samples, five models, four languages, eight prompt types, two variants per combination, roughly 100,800 words total, every sample generated with identical instructions.
The original analysis averaged across all five models to build the "Average AI" locale baselines. This post keeps the same underlying data and pulls the models apart instead, isolating the English-language corpus (80 samples: 5 models × 8 prompt types × 2 variants) to compare model against model on equal footing.
Every number in this post is the English slice unless labeled otherwise. That is deliberate, not a shortcut. The scoring formulas are language-aware: sentence segmentation, tokenization, and word lists all switch on locale, so scores are only comparable to other scores computed in the same language. A single cross-language "per-model" average would be mixing measurements taken with different instruments. Where a cross-check is useful, we quote the French and Spanish slices side by side rather than blending them.
The six axes are unchanged: sentence complexity, vocabulary range, expressiveness, formality, consistency, conciseness. Same deterministic, language-aware formulas described in the methodology post.
The Per-Model Numbers (English Corpus)
| Axis | Opus 4.6 | Sonnet 4.6 | Haiku 4.5 | GPT-5.2 | Gemini 3 Pro | Spread |
|---|---|---|---|---|---|---|
| Sentence Complexity | 60 | 79 | 69 | 72 | 50 | 29 |
| Vocabulary Range | 48 | 53 | 53 | 50 | 40 | 13 |
| Expressiveness | 100 | 84 | 84 | 65 | 39 | 61 |
| Formality | 51 | 45 | 70 | 65 | 64 | 25 |
| Consistency | 53 | 47 | 54 | 53 | 63 | 16 |
| Conciseness | 48 | 32 | 35 | 33 | 52 | 20 |
Two rows barely move. One row moves more than any other by a wide margin. The rest sit in between.
Where Models Actually Agree
Vocabulary range is the tightest cluster in the English slice: a 13-point spread (40-53) across all five models. Type-Token Ratio is remarkably stable regardless of which model generated the text. If you're trying to get AI to use richer or more constrained vocabulary, switching models is close to a wash. The lever that moves this axis is your Style Profile, not model choice.
Consistency is nearly as tight: a 16-point spread (47-63). Sentence-length rhythm is one of the more model-agnostic properties of AI writing. This lines up with the broader pattern we describe in the median user problem: RLHF pushes every model toward a "safe," non-bursty cadence, and that pressure seems to land on consistency almost identically across labs.
On these two axes, "Average AI" is a genuinely meaningful abstraction. Calibrating your Style Profile's override against the multi-model baseline works no matter which model reads your Master Prompt.
Where Models Pull Apart
Expressiveness is the outlier, and it isn't close. Claude Opus 4.6 sits at 100, driven by heavier use of exclamations, questions, attitude markers, and em-dashes. Gemini 3 Pro sits at 39. That 61-point gap is more than double the next-widest spread in the English slice. If you're targeting Opus specifically, the "sound less like AI" override you need on this axis is substantially more aggressive than if you're targeting Gemini.
Two caveats on that 61, both of which cut in the same direction. First, 100 is our scale ceiling, not Opus's measured peak: the raw score is clamped, so the true distance between Opus and Gemini is at least 61 points and we can't say how much more. Second, a one-language result invites the obvious question of whether it's an artifact of English. It isn't. The same top-to-bottom expressiveness spread appears independently in the other two languages we can score reliably:
| Language | Most expressive | Least expressive | Spread |
|---|---|---|---|
| English | Opus 4.6 (100) | Gemini 3 Pro (39) | 61 |
| French | Opus 4.6 (95) | Gemini 3 Pro (39) | 56 |
| Spanish | Haiku 4.5 (88) | Gemini 3 Pro (28) | 60 |
Gemini 3 Pro anchors the bottom in all three, and a Claude model anchors the top in all three. That consistency is the reason we're willing to call the expressiveness divergence real rather than a quirk of one corpus cut. (Our Japanese slice is excluded here; see the limitations note below for why it can't currently support this comparison.)
Sentence complexity shows a real but smaller 29-point spread. Claude Sonnet 4.6 produces the most structurally dense sentences (79); Gemini 3 Pro produces the plainest (50). This one runs counter to the "tight convergence" story you might expect from the locale-level averages, which blend all five models together and smooth this variance out.
Formality splits by 25 points, and the direction is counterintuitive. Claude Haiku 4.5, the smallest and fastest model in the corpus, scores highest (70). Claude Sonnet 4.6 scores lowest (45). GPT-5.2 (65) and Gemini 3 Pro (64) land close to Haiku, not close to Sonnet. If your mental model is "smaller models are more casual," this data says otherwise, at least for this model family lineup at this snapshot in time.
Conciseness spreads 20 points. Claude Sonnet 4.6 produces the tightest sentences after the inverse-length transform (32); Gemini 3 Pro produces the loosest (52). Sonnet's combination of high sentence complexity (79) and low conciseness (32) is its own signature: long, structurally elaborate sentences that don't compress well.
Why This Matters for Your Style Profile
If you're building a Style Profile, you're implicitly choosing between two calibration strategies: override against the multi-model Average AI baseline, or override against the specific model you use most.
The multi-model baseline is the right default. It's what powers your Writing DNA Snapshot, and it produces an override that transfers across ChatGPT, Claude, and Gemini without re-tuning. For vocabulary range and consistency, this is close to free: the models barely disagree, so a multi-model target is nearly identical to any single-model target.
Expressiveness is the exception worth knowing about. Because the spread is at least 61 points in English, and 56-60 in French and Spanish, wide enough that "average" and "closest individual model" can disagree by a large margin, a multi-model target can undershoot how far you actually need to push against a specific model's defaults. If you know you're writing primarily inside Claude Opus 4.6, treat expressiveness as the one axis where a model-specific override, not just the Average AI delta, is worth the extra precision. For every other axis, the multi-model baseline your Snapshot already gives you is doing its job.
What the Average Was Hiding
The "Average AI" baseline is still the right default target: it's what makes a Style Profile portable across platforms instead of tuned to one model's personality. But averaging five models into one shape hides that the models don't actually agree with each other on how expressive AI writing should be. They agree closely on vocabulary and rhythm. They disagree by at least 61 points, in English, on how much rhetorical energy to put on the page.
If you've felt that Claude "sounds different" from ChatGPT or Gemini, the data backs that up, and pinpoints exactly which axis carries the difference.
For the full baseline methodology and per-locale tables, see how we measure average AI. For the training-level explanation of why models converge at all, see the median user problem.
Limitations
Three things worth knowing before you lean on these numbers.
The measurements are a February 2026 snapshot. The corpus was generated on 2026-02-15 against Claude Opus 4.6, Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.2, and Gemini 3 Pro. Every one of those model versions has since been superseded. Model defaults move with each release, so treat the per-model scores as a record of how that lineup behaved at that moment, not as current behavior. The structural finding we'd expect to survive is that models differ far more on expressiveness than on vocabulary or rhythm; the specific point values almost certainly do not.
Expressiveness is capped at 100, and Opus is sitting on the cap. Any score at the ceiling is a lower bound. Where we quote a 61-point spread, read it as "at least 61."
The Japanese slice is excluded from every comparison in this post. In our Japanese samples, four of the five models (Opus, Sonnet, Haiku, and Gemini) all score an identical, perfect 100 on expressiveness. Four models from two different labs landing on exactly the ceiling is a measurement artifact, not a finding. The cause is in our scoring code, not in the writing: the expressiveness formula divides several of its punctuation terms by a word count produced by an ASCII-only tokenizer, which massively undercounts Japanese text and inflates those terms until the score saturates. Japanese expressiveness needs to be re-validated against a locale-aware word count before it can be cited as a comparative result — including the Japanese column in our earlier methodology post, which we are revisiting. The other five axes and the other three languages are unaffected by this specific issue.
Ready to See Where You Fall?
Want your own six-axis radar chart against the multi-model Average AI baseline, including your delta on the axis that matters most for whichever model you use?
Get Your Free Writing DNA Snapshot
Submit a few writing samples. See your scores on all six axes, and exactly where your writing diverges from Average AI. Free, no credit card.
Your writing has a fingerprint. We measure it. My Writing Twin turns that measurement into instructions that make any AI (ChatGPT, Claude, Gemini) write like you.
Make AI write like you, not like a bot
You just read how to tune AI output by hand. MyWritingTwin does it from your real writing: paste a few samples, get a voice profile that works in ChatGPT, Claude, and Gemini.
Create your free writing profileNot ready yet? Get our guide to AI voice profiles by email instead.