IntellectuaLead
English

Independent style analysis

Stylistic LLM Writing benchmark

Compare frontier language models writing quality based on stylistic metrics (grammatical, syntactic, rethorical, lexical, prosody...) and with human writing as a reference. Lower AI evidence means the model output is stylistically closer to the human baseline used in this benchmark. Since the leaderboard is based on a stylistic analysis and not blind human preferences or machine-learning discrimination like AI detectors, you can understand why a model sounds more human than another.

41 compared models and baseline 8 writing genres 9 equally weighted signals

Detailed model comparison

Scores update automatically when a new benchmark is published.

Low AI Medium AI High AI
RankModelLatest published runsAI ScoreAverage of nine signalsSentence length diversityDistribution breadth, quantiles and tailsGPT-2 perplexity similarityModel-like token predictabilityNoun contextualizationContext supplied around nounsHeavy grammarLLM-favored constructions per 1,000 lexical tokensPersonal tonePerspective signals per 100 wordsAI SlopRhetorical templates per 1,000 wordsAI WordingOverused wording per 1,000 wordsPunctuation richnessRare-sensitive variety across ten mark familiesSection variationIntro-body-ending change
#117.5/100Low AICombined stylistic evidence is low.0.741Low AISentence lengths are broadly distributed.0.0056Low AIToken choices are within the lowest evidence quarter.0.81Low AINouns are well contextualized with linking, helping, and qualifying language.23/1kLow AIFew of the directional heavy-grammar signals are elevated.7.4%Low AIPersonal-tone density is near the human anchor.7.8/1kLow AIFew repeated rhetorical templates are present.5.2/1kLow AIFew overused wording matches are present.0.767Low AIPunctuation is richly distributed across common and less-common mark types.0.296Low AIThe document changes shape naturally.
#234.3/100Low AICombined stylistic evidence is low.0.672High AISentence lengths are concentrated in a few bands.0.001Low AIToken choices are within the lowest evidence quarter.0.567High AINouns receive little relational or qualifying context.21.1/1kLow AIFew of the directional heavy-grammar signals are elevated.6.1%Medium AISome personal or subjective language is visible.12.8/1kMedium AISome repeated rhetorical templates are present.7.5/1kLow AIFew overused wording matches are present.0.758Low AIPunctuation is richly distributed across common and less-common mark types.0.303Low AIThe document changes shape naturally.
#336.8/100Medium AICombined stylistic evidence is moderate.0.755Low AISentence lengths are broadly distributed.0Low AIToken choices are within the lowest evidence quarter.0.747Low AINouns are well contextualized with linking, helping, and qualifying language.23.6/1kLow AIFew of the directional heavy-grammar signals are elevated.5.1%High AIVery little perspective or uncertainty is visible.15/1kHigh AIRepeated rhetorical templates are frequent.4.9/1kLow AIFew overused wording matches are present.0.596High AIPunctuation is concentrated in relatively few mark types.0.209High AIVery little style change appears across the document.
#438.2/100Medium AICombined stylistic evidence is moderate.0.604High AISentence lengths are concentrated in a few bands.0.0004Low AIToken choices are within the lowest evidence quarter.0.648High AINouns receive little relational or qualifying context.26/1kLow AIFew of the directional heavy-grammar signals are elevated.5.9%High AIVery little perspective or uncertainty is visible.8.5/1kLow AIFew repeated rhetorical templates are present.7.8/1kLow AIFew overused wording matches are present.0.646Medium AISome less-common punctuation types are represented.0.218High AIVery little style change appears across the document.
#538.4/100Medium AICombined stylistic evidence is moderate.0.701Medium AISentence lengths use several bands.0.0001Low AIToken choices are within the lowest evidence quarter.0.717Medium AINoun contextualization is moderate.23.3/1kLow AIFew of the directional heavy-grammar signals are elevated.5.3%High AIVery little perspective or uncertainty is visible.15.9/1kHigh AIRepeated rhetorical templates are frequent.5.1/1kLow AIFew overused wording matches are present.0.609High AIPunctuation is concentrated in relatively few mark types.0.249Medium AISome section-to-section change is visible.
#638.4/100Medium AICombined stylistic evidence is moderate.0.616High AISentence lengths are concentrated in a few bands.0.0157Low AIToken choices are within the lowest evidence quarter.0.706Medium AINoun contextualization is moderate.21.2/1kLow AIFew of the directional heavy-grammar signals are elevated.6.6%Medium AISome personal or subjective language is visible.10.6/1kLow AIFew repeated rhetorical templates are present.8.2/1kLow AIFew overused wording matches are present.0.605High AIPunctuation is concentrated in relatively few mark types.0.217High AIVery little style change appears across the document.
#738.9/100Medium AICombined stylistic evidence is moderate.0.727Low AISentence lengths are broadly distributed.0.0897Medium AIToken choices are between one quarter and one half of the evidence range.0.711Medium AINoun contextualization is moderate.28.4/1kLow AIFew of the directional heavy-grammar signals are elevated.5%High AIVery little perspective or uncertainty is visible.11.3/1kMedium AISome repeated rhetorical templates are present.8.4/1kLow AIFew overused wording matches are present.0.696Low AIPunctuation is richly distributed across common and less-common mark types.0.182High AIVery little style change appears across the document.
#840.5/100Medium AICombined stylistic evidence is moderate.0.673High AISentence lengths are concentrated in a few bands.0.0284Low AIToken choices are within the lowest evidence quarter.0.668Medium AINoun contextualization is moderate.28.3/1kLow AIFew of the directional heavy-grammar signals are elevated.5.6%High AIVery little perspective or uncertainty is visible.15.4/1kHigh AIRepeated rhetorical templates are frequent.10.6/1kMedium AISome overused wording matches are present.0.684Medium AISome less-common punctuation types are represented.0.244Medium AISome section-to-section change is visible.
#940.5/100Medium AICombined stylistic evidence is moderate.0.707Medium AISentence lengths use several bands.0.0004Low AIToken choices are within the lowest evidence quarter.0.792Low AINouns are well contextualized with linking, helping, and qualifying language.29.6/1kLow AIFew of the directional heavy-grammar signals are elevated.6.2%Medium AISome personal or subjective language is visible.16/1kHigh AIRepeated rhetorical templates are frequent.8.2/1kLow AIFew overused wording matches are present.0.531High AIPunctuation is concentrated in relatively few mark types.0.235Medium AISome section-to-section change is visible.
#1041.1/100Medium AICombined stylistic evidence is moderate.0.672High AISentence lengths are concentrated in a few bands.0.0145Low AIToken choices are within the lowest evidence quarter.0.745Low AINouns are well contextualized with linking, helping, and qualifying language.24.5/1kLow AIFew of the directional heavy-grammar signals are elevated.6.5%Medium AISome personal or subjective language is visible.15.5/1kHigh AIRepeated rhetorical templates are frequent.8.5/1kLow AIFew overused wording matches are present.0.58High AIPunctuation is concentrated in relatively few mark types.0.213High AIVery little style change appears across the document.
#1141.1/100Medium AICombined stylistic evidence is moderate.0.709Medium AISentence lengths use several bands.0.0087Low AIToken choices are within the lowest evidence quarter.0.457High AINouns receive little relational or qualifying context.32.1/1kLow AIFew of the directional heavy-grammar signals are elevated.4.6%High AIVery little perspective or uncertainty is visible.10.5/1kLow AIFew repeated rhetorical templates are present.18/1kHigh AIOverused wording matches are frequent.0.733Low AIPunctuation is richly distributed across common and less-common mark types.0.247Medium AISome section-to-section change is visible.
#1241.5/100Medium AICombined stylistic evidence is moderate.0.66High AISentence lengths are concentrated in a few bands.0.036Low AIToken choices are within the lowest evidence quarter.0.709Medium AINoun contextualization is moderate.29.6/1kLow AIFew of the directional heavy-grammar signals are elevated.6.2%Medium AISome personal or subjective language is visible.10/1kLow AIFew repeated rhetorical templates are present.10.9/1kMedium AISome overused wording matches are present.0.631Medium AISome less-common punctuation types are represented.0.203High AIVery little style change appears across the document.
#1341.6/100Medium AICombined stylistic evidence is moderate.0.667High AISentence lengths are concentrated in a few bands.0.0033Low AIToken choices are within the lowest evidence quarter.0.772Low AINouns are well contextualized with linking, helping, and qualifying language.28.7/1kLow AIFew of the directional heavy-grammar signals are elevated.5.8%High AIVery little perspective or uncertainty is visible.12.6/1kMedium AISome repeated rhetorical templates are present.9.7/1kMedium AISome overused wording matches are present.0.56High AIPunctuation is concentrated in relatively few mark types.0.199High AIVery little style change appears across the document.
#1441.8/100Medium AICombined stylistic evidence is moderate.0.714Medium AISentence lengths use several bands.0.0011Low AIToken choices are within the lowest evidence quarter.0.612High AINouns receive little relational or qualifying context.28.4/1kLow AIFew of the directional heavy-grammar signals are elevated.5.6%High AIVery little perspective or uncertainty is visible.17.4/1kHigh AIRepeated rhetorical templates are frequent.10.8/1kMedium AISome overused wording matches are present.0.731Low AIPunctuation is richly distributed across common and less-common mark types.0.214High AIVery little style change appears across the document.
#1542.1/100Medium AICombined stylistic evidence is moderate.0.683High AISentence lengths are concentrated in a few bands.0Low AIToken choices are within the lowest evidence quarter.0.68Medium AINoun contextualization is moderate.23.7/1kLow AIFew of the directional heavy-grammar signals are elevated.5.6%High AIVery little perspective or uncertainty is visible.17.3/1kHigh AIRepeated rhetorical templates are frequent.8/1kLow AIFew overused wording matches are present.0.634Medium AISome less-common punctuation types are represented.0.23Medium AISome section-to-section change is visible.
#1642.7/100Medium AICombined stylistic evidence is moderate.0.695High AISentence lengths are concentrated in a few bands.0.0021Low AIToken choices are within the lowest evidence quarter.0.536High AINouns receive little relational or qualifying context.25.9/1kLow AIFew of the directional heavy-grammar signals are elevated.5.8%High AIVery little perspective or uncertainty is visible.13.9/1kMedium AISome repeated rhetorical templates are present.10.7/1kMedium AISome overused wording matches are present.0.682Medium AISome less-common punctuation types are represented.0.23Medium AISome section-to-section change is visible.
#1743.8/100Medium AICombined stylistic evidence is moderate.0.743Low AISentence lengths are broadly distributed.0.0104Low AIToken choices are within the lowest evidence quarter.0.507High AINouns receive little relational or qualifying context.34.2/1kLow AIFew of the directional heavy-grammar signals are elevated.4.1%High AIVery little perspective or uncertainty is visible.11/1kLow AIFew repeated rhetorical templates are present.14.2/1kHigh AIOverused wording matches are frequent.0.663Medium AISome less-common punctuation types are represented.0.272Low AIThe document changes shape naturally.
#1843.9/100Medium AICombined stylistic evidence is moderate.0.699Medium AISentence lengths are concentrated in a few bands.0.1031High AIToken choices are beyond the evidence-range midpoint.0.692Medium AINoun contextualization is moderate.32.4/1kLow AIFew of the directional heavy-grammar signals are elevated.4.7%High AIVery little perspective or uncertainty is visible.10.5/1kLow AIFew repeated rhetorical templates are present.20.9/1kHigh AIOverused wording matches are frequent.0.692Low AIPunctuation is richly distributed across common and less-common mark types.0.181High AIVery little style change appears across the document.
#1944.5/100Medium AICombined stylistic evidence is moderate.0.672High AISentence lengths are concentrated in a few bands.0.9996High AIToken choices are beyond the evidence-range midpoint.0.663Medium AINoun contextualization is moderate.25.3/1kLow AIFew of the directional heavy-grammar signals are elevated.6%Medium AISome personal or subjective language is visible.5.7/1kLow AIFew repeated rhetorical templates are present.15.6/1kHigh AIOverused wording matches are frequent.0.688Medium AISome less-common punctuation types are represented.0.293Low AIThe document changes shape naturally.
#2044.8/100Medium AICombined stylistic evidence is moderate.0.626High AISentence lengths are concentrated in a few bands.0.0295Low AIToken choices are within the lowest evidence quarter.0.729Low AINouns are well contextualized with linking, helping, and qualifying language.25.3/1kLow AIFew of the directional heavy-grammar signals are elevated.6.4%Medium AISome personal or subjective language is visible.10.9/1kLow AIFew repeated rhetorical templates are present.7.7/1kLow AIFew overused wording matches are present.0.566High AIPunctuation is concentrated in relatively few mark types.0.159High AIVery little style change appears across the document.
#2146.1/100Medium AICombined stylistic evidence is moderate.0.676High AISentence lengths are concentrated in a few bands.0.0002Low AIToken choices are within the lowest evidence quarter.0.525High AINouns receive little relational or qualifying context.34/1kLow AIFew of the directional heavy-grammar signals are elevated.5%High AIVery little perspective or uncertainty is visible.8.1/1kLow AIFew repeated rhetorical templates are present.27.9/1kHigh AIOverused wording matches are frequent.0.726Low AIPunctuation is richly distributed across common and less-common mark types.0.259Medium AISome section-to-section change is visible.
#2246.9/100Medium AICombined stylistic evidence is moderate.0.636High AISentence lengths are concentrated in a few bands.0.007Low AIToken choices are within the lowest evidence quarter.0.619High AINouns receive little relational or qualifying context.33.4/1kLow AIFew of the directional heavy-grammar signals are elevated.6.8%Low AIPersonal-tone density is near the human anchor.15/1kHigh AIRepeated rhetorical templates are frequent.15.7/1kHigh AIOverused wording matches are frequent.0.651Medium AISome less-common punctuation types are represented.0.197High AIVery little style change appears across the document.
#2346.9/100Medium AICombined stylistic evidence is moderate.0.654High AISentence lengths are concentrated in a few bands.0.2912High AIToken choices are beyond the evidence-range midpoint.0.652Medium AINoun contextualization is moderate.28.7/1kLow AIFew of the directional heavy-grammar signals are elevated.5.2%High AIVery little perspective or uncertainty is visible.10.8/1kLow AIFew repeated rhetorical templates are present.20.9/1kHigh AIOverused wording matches are frequent.0.702Low AIPunctuation is richly distributed across common and less-common mark types.0.216High AIVery little style change appears across the document.
#2446.9/100Medium AICombined stylistic evidence is moderate.0.579High AISentence lengths are concentrated in a few bands.0.0051Low AIToken choices are within the lowest evidence quarter.0.6High AINouns receive little relational or qualifying context.23.9/1kLow AIFew of the directional heavy-grammar signals are elevated.7.1%Low AIPersonal-tone density is near the human anchor.12.2/1kMedium AISome repeated rhetorical templates are present.6.9/1kLow AIFew overused wording matches are present.0.625Medium AISome less-common punctuation types are represented.0.213High AIVery little style change appears across the document.
#2546.9/100Medium AICombined stylistic evidence is moderate.0.669High AISentence lengths are concentrated in a few bands.0.0116Low AIToken choices are within the lowest evidence quarter.0.588High AINouns receive little relational or qualifying context.30/1kLow AIFew of the directional heavy-grammar signals are elevated.5%High AIVery little perspective or uncertainty is visible.10.4/1kLow AIFew repeated rhetorical templates are present.10.5/1kMedium AISome overused wording matches are present.0.563High AIPunctuation is concentrated in relatively few mark types.0.194High AIVery little style change appears across the document.
#2647.1/100Medium AICombined stylistic evidence is moderate.0.677High AISentence lengths are concentrated in a few bands.0.629High AIToken choices are beyond the evidence-range midpoint.0.747Low AINouns are well contextualized with linking, helping, and qualifying language.30.1/1kLow AIFew of the directional heavy-grammar signals are elevated.5.5%High AIVery little perspective or uncertainty is visible.11/1kLow AIFew repeated rhetorical templates are present.17.5/1kHigh AIOverused wording matches are frequent.0.694Low AIPunctuation is richly distributed across common and less-common mark types.0.209High AIVery little style change appears across the document.
#2747.2/100Medium AICombined stylistic evidence is moderate.0.662High AISentence lengths are concentrated in a few bands.0.02Low AIToken choices are within the lowest evidence quarter.0.677Medium AINoun contextualization is moderate.31.3/1kLow AIFew of the directional heavy-grammar signals are elevated.5.3%High AIVery little perspective or uncertainty is visible.13.8/1kMedium AISome repeated rhetorical templates are present.10.8/1kMedium AISome overused wording matches are present.0.601High AIPunctuation is concentrated in relatively few mark types.0.175High AIVery little style change appears across the document.
#2848.2/100Medium AICombined stylistic evidence is moderate.0.658High AISentence lengths are concentrated in a few bands.0.0028Low AIToken choices are within the lowest evidence quarter.0.573High AINouns receive little relational or qualifying context.32.7/1kLow AIFew of the directional heavy-grammar signals are elevated.5.1%High AIVery little perspective or uncertainty is visible.13.3/1kMedium AISome repeated rhetorical templates are present.8.8/1kLow AIFew overused wording matches are present.0.659Medium AISome less-common punctuation types are represented.0.189High AIVery little style change appears across the document.
#2948.5/100Medium AICombined stylistic evidence is moderate.0.652High AISentence lengths are concentrated in a few bands.0.349High AIToken choices are beyond the evidence-range midpoint.0.73Low AINouns are well contextualized with linking, helping, and qualifying language.30.2/1kLow AIFew of the directional heavy-grammar signals are elevated.6.3%Medium AISome personal or subjective language is visible.11.6/1kMedium AISome repeated rhetorical templates are present.20/1kHigh AIOverused wording matches are frequent.0.647Medium AISome less-common punctuation types are represented.0.182High AIVery little style change appears across the document.
#3048.7/100Medium AICombined stylistic evidence is moderate.0.685High AISentence lengths are concentrated in a few bands.0.3445High AIToken choices are beyond the evidence-range midpoint.0.73Low AINouns are well contextualized with linking, helping, and qualifying language.31.6/1kLow AIFew of the directional heavy-grammar signals are elevated.5.5%High AIVery little perspective or uncertainty is visible.12.4/1kMedium AISome repeated rhetorical templates are present.21.1/1kHigh AIOverused wording matches are frequent.0.637Medium AISome less-common punctuation types are represented.0.217High AIVery little style change appears across the document.
#3148.8/100Medium AICombined stylistic evidence is moderate.0.615High AISentence lengths are concentrated in a few bands.0.0051Low AIToken choices are within the lowest evidence quarter.0.624High AINouns receive little relational or qualifying context.35.9/1kLow AIFew of the directional heavy-grammar signals are elevated.6.6%Medium AISome personal or subjective language is visible.12.7/1kMedium AISome repeated rhetorical templates are present.13.9/1kHigh AIOverused wording matches are frequent.0.636Medium AISome less-common punctuation types are represented.0.215High AIVery little style change appears across the document.
#3249.5/100Medium AICombined stylistic evidence is moderate.0.722Medium AISentence lengths use several bands.0.6349High AIToken choices are beyond the evidence-range midpoint.0.73Low AINouns are well contextualized with linking, helping, and qualifying language.28.1/1kLow AIFew of the directional heavy-grammar signals are elevated.5.5%High AIVery little perspective or uncertainty is visible.16.7/1kHigh AIRepeated rhetorical templates are frequent.15.3/1kHigh AIOverused wording matches are frequent.0.703Low AIPunctuation is richly distributed across common and less-common mark types.0.187High AIVery little style change appears across the document.
#3350.1/100High AICombined stylistic evidence is high.0.669High AISentence lengths are concentrated in a few bands.0Low AIToken choices are within the lowest evidence quarter.0.455High AINouns receive little relational or qualifying context.34.3/1kLow AIFew of the directional heavy-grammar signals are elevated.4.6%High AIVery little perspective or uncertainty is visible.10.2/1kLow AIFew repeated rhetorical templates are present.19.7/1kHigh AIOverused wording matches are frequent.0.654Medium AISome less-common punctuation types are represented.0.226High AIVery little style change appears across the document.
#3450.4/100High AICombined stylistic evidence is high.0.661High AISentence lengths are concentrated in a few bands.0.1528High AIToken choices are beyond the evidence-range midpoint.0.67Medium AINoun contextualization is moderate.30.9/1kLow AIFew of the directional heavy-grammar signals are elevated.5.6%High AIVery little perspective or uncertainty is visible.10/1kLow AIFew repeated rhetorical templates are present.22.9/1kHigh AIOverused wording matches are frequent.0.609High AIPunctuation is concentrated in relatively few mark types.0.236Medium AISome section-to-section change is visible.
#3550.5/100High AICombined stylistic evidence is high.0.706Medium AISentence lengths use several bands.0.0128Low AIToken choices are within the lowest evidence quarter.0.562High AINouns receive little relational or qualifying context.42.2/1kMedium AISome dense constructions are elevated.5.6%High AIVery little perspective or uncertainty is visible.14.7/1kHigh AIRepeated rhetorical templates are frequent.18.3/1kHigh AIOverused wording matches are frequent.0.609High AIPunctuation is concentrated in relatively few mark types.0.191High AIVery little style change appears across the document.
#3651.1/100High AICombined stylistic evidence is high.0.679High AISentence lengths are concentrated in a few bands.0.2718High AIToken choices are beyond the evidence-range midpoint.0.691Medium AINoun contextualization is moderate.28.3/1kLow AIFew of the directional heavy-grammar signals are elevated.5.8%High AIVery little perspective or uncertainty is visible.11.6/1kMedium AISome repeated rhetorical templates are present.26.8/1kHigh AIOverused wording matches are frequent.0.726Low AIPunctuation is richly distributed across common and less-common mark types.0.152High AIVery little style change appears across the document.
#3751.3/100High AICombined stylistic evidence is high.0.647High AISentence lengths are concentrated in a few bands.0.0007Low AIToken choices are within the lowest evidence quarter.0.583High AINouns receive little relational or qualifying context.34.6/1kLow AIFew of the directional heavy-grammar signals are elevated.5%High AIVery little perspective or uncertainty is visible.20/1kHigh AIRepeated rhetorical templates are frequent.23.5/1kHigh AIOverused wording matches are frequent.0.757Low AIPunctuation is richly distributed across common and less-common mark types.0.21High AIVery little style change appears across the document.
#3853.2/100High AICombined stylistic evidence is high.0.569High AISentence lengths are concentrated in a few bands.0.024Low AIToken choices are within the lowest evidence quarter.0.593High AINouns receive little relational or qualifying context.25.2/1kLow AIFew of the directional heavy-grammar signals are elevated.6.9%Low AIPersonal-tone density is near the human anchor.12.1/1kMedium AISome repeated rhetorical templates are present.9.7/1kMedium AISome overused wording matches are present.0.521High AIPunctuation is concentrated in relatively few mark types.0.152High AIVery little style change appears across the document.
#3954.8/100High AICombined stylistic evidence is high.0.642High AISentence lengths are concentrated in a few bands.0.0001Low AIToken choices are within the lowest evidence quarter.0.563High AINouns receive little relational or qualifying context.47.5/1kMedium AISome dense constructions are elevated.6.3%Medium AISome personal or subjective language is visible.16.7/1kHigh AIRepeated rhetorical templates are frequent.14.2/1kHigh AIOverused wording matches are frequent.0.619Medium AISome less-common punctuation types are represented.0.22High AIVery little style change appears across the document.
#4056.4/100High AICombined stylistic evidence is high.0.663High AISentence lengths are concentrated in a few bands.0.9944High AIToken choices are beyond the evidence-range midpoint.0.714Medium AINoun contextualization is moderate.30.1/1kLow AIFew of the directional heavy-grammar signals are elevated.6.7%Low AIPersonal-tone density is near the human anchor.4.8/1kLow AIFew repeated rhetorical templates are present.16.4/1kHigh AIOverused wording matches are frequent.0.441High AIPunctuation is concentrated in relatively few mark types.0.213High AIVery little style change appears across the document.
#4158.8/100High AICombined stylistic evidence is high.0.61High AISentence lengths are concentrated in a few bands.0.4131High AIToken choices are beyond the evidence-range midpoint.0.54High AINouns receive little relational or qualifying context.27.8/1kLow AIFew of the directional heavy-grammar signals are elevated.5.6%High AIVery little perspective or uncertainty is visible.12.7/1kMedium AISome repeated rhetorical templates are present.26.4/1kHigh AIOverused wording matches are frequent.0.645Medium AISome less-common punctuation types are represented.0.266Low AIThe document changes shape naturally.
Labels summarize low, medium or high stylistic evidence. They compare writing tendencies and do not prove authorship.

Benchmark methodology

Nine writing signals used in every AI Score

Every model is evaluated with the same nine stylistic signals and the same set of writing genres. The score summarizes those signals across each model's completed tasks.

Sentence length diversity

Looks at whether a text naturally mixes short, medium and long sentences, including the center and tails of the distribution. Broader, less repetitive pacing is treated as more human-like.

GPT-2 perplexity similarity

Compares token predictability with patterns associated with language-model writing. It is one probabilistic signal, not a standalone detector.

Noun contextualization

Divides adpositions, auxiliaries and adverbs by nouns. A higher value means noun-led ideas receive more relational, tense or qualifying context. The calibration gives no AI evidence at 0.80 and full evidence at 0.50.

Heavy grammar

Tracks dense constructions including participial clauses, that-subject clauses, and 2-4 word adjective/noun sequences ending in an abstract -ty/-ity, -tion/-sion, -ness, or -ment noun.

Personal tone

Looks for four stance families: self-mentions including I, we, you, and your; hedges; boosters; and attitude markers, plus direct questions. Modal auxiliaries and multiword phrases such as “in practice” or “in fact” are classified inside those families and are not counted twice. A value of 7.4% gives zero AI evidence and 4.6% gives full AI evidence.

AI Slop

Tracks repeated rhetorical templates such as em dashes, three-part sequences, formulaic contrasts and while-led concessions. Occasional use is normal; repetition is more informative.

AI Wording

Tracks words and short phrases that appear unusually often in LLM output. The metric is most useful when several generic matches accumulate in one text.

Section variation

Checks whether the opening, body and ending perform differently or repeat the same stylistic shape. More purposeful change is treated as more human-like.

Punctuation richness

Uses rare-sensitive Hill diversity at q = 0.25 across ten punctuation families. It rewards meaningful representation of less-common marks more strongly than ordinary entropy and contributes equally to the AI Score.

Controlled comparison

Eight writing genres

Using different genres reduces the chance that one topic or writing format determines the leaderboard.

  1. SEO article about coffee roasting
  2. Academic essay on universal basic income
  3. Research writing about AI detection
  4. Personal writing about overcoming insecurity
  5. Fiction about a Halloween party going wrong
  6. Biography of Alan Turing
  7. Copywriting email for a dating service
  8. News report on New York's cost of living

Questions

About the Stylistic LLM Writing benchmark

What does the AI Score measure?

It summarizes eight interpretable stylistic signals across the same controlled writing tasks. A lower result means the model's outputs are stylistically closer to the human reference used here; it does not mean every individual passage will sound human.

Does a high score prove that a text was written by AI?

No. The score compares measurable stylistic tendencies. It is useful for controlled model comparison, but it is not proof of authorship for an isolated document.

Why is the human baseline included?

The human samples provide a consistent reference under the same eight genres and tasks. They make the direction and scale of each model result easier to interpret without claiming that one human style is universally ideal.

How is this different from a preference leaderboard?

Preference leaderboards report which answer readers or judges favor. This benchmark measures explainable properties of the writing itself, including grammar, syntax, rhetorical habits, wording, sentence pacing and section structure.

How is this different from an AI detector?

AI detectors usually estimate whether an isolated text belongs to an AI or human class. This benchmark compares models under controlled prompts and exposes the stylistic signals behind the ranking. It should not be used to accuse an author or verify authorship.

Why are several genres used?

Writing style changes with purpose. Testing SEO, academic, research, personal, fiction, biography, email and news writing reduces the chance that one topic or format decides the ranking.

Why can a model's ranking change?

A ranking may change after new task outputs are published, a provider updates a model, missing analyses are completed, or the benchmark's linguistic analysis is improved. The table always uses the latest published runs.

Can two models with similar AI Scores still write differently?

Yes. Their overall scores can be close while their metric profiles differ substantially. Open the model details to compare which model relies more on dense grammar, formulaic rhetoric, predictable wording or uniform section structure.

How often does the leaderboard update?

The website reads the latest benchmark published from the Intellectualead app. Completed NLP backfills and new benchmark runs publish refreshed metrics automatically.