Dataset

AI writing baselines by language (320 samples, 5 models, 4 locales)

This dataset gives the average stylometric profile of AI-generated text per language, measured from 320 samples across 5 models, 4 locales and 8 prompt types, and it is the baseline a human writer's profile is compared against.

Saying AI writing sounds generic is an assertion until there is a number for generic. These are those numbers: the centroid of AI output per language on six stylometric dimensions, measured rather than estimated, so a writer's own scores can be positioned against it.

The headline result is that the baseline is not the same in every language. Conciseness in French sits far below English, and the Japanese expressiveness figure hits the ceiling of the formula because Japanese business text uses question forms and polite markers the measure counts as expressive. A tool comparing a French writer against an English baseline gets the answer wrong by construction.

DimensionEnglishFrenchSpanishJapanese
Sentence complexity65757162
Vocabulary richness48494437
Expressiveness767459100
Formality58424659
Consistency53525553
Conciseness42323645
Average AI-generated text scores (0-100) by language, measured 2026-02-15

Method notes

  • Sample design: 5 models (Opus 4.6, Sonnet 4.5, Haiku 4.5, GPT-5.2, Gemini 3 Pro) x 4 locales x 8 prompt types x 2 variants = 320 samples.
  • Japanese expressiveness reaches 100 because Japanese business text uses many question forms and polite markers the expressiveness formula counts. AI baselines and user scores use the same formula, so they stay comparable.
  • German is not in this table: it is a supported locale for word lists but has no measured baseline, so a German comparison would be against the wrong centroid.
  • Scores are 0-100 on MyWritingTwin's own formulas and are not comparable to any other tool's dimension of the same name.

Frequently asked questions

When were these AI baselines measured?
2026-02-15, from 320 samples across five models, four locales and eight prompt types. The measurement date is published because model behaviour changes and a baseline without a date is not a usable figure.
Why is there no German baseline?
Because it has not been measured. German is supported for word-list analysis, but comparing a German writer against the English centroid would produce a confident wrong answer, so the baseline is reported as absent rather than substituted.
What does a high or low score mean?
Each dimension is 0-100 on MyWritingTwin's formulas. The useful reading is not the absolute value but the gap between a writer's score and the baseline for their language, which is where the voice actually lives.

Read further

Sources and how to cite this page

Cite this page

MyWritingTwin. “AI writing baselines by language (320 samples, 5 models, 4 locales).” Last updated 2026-09-15. https://www.mywritingtwin.com/reference/ai-baseline-divergence