Compare how frontier language models write across the same set of practical tasks. Lower AI evidence means the model output is stylistically closer to the human baseline used in this benchmark.
25 compared models and baseline8 writing genres6 equally weighted signals
Detailed model comparison
Scores update automatically when a new benchmark is published.
Low AIMedium AIHigh AI
Rank
ModelLatest published runs
AI ScoreAverage of six signals
Sentence rhythmSentence-length variation
GPT-2 proximityModel-like predictability
Syntax diversityGrammar-pattern variety
Prose heavinessDense constructions per 1,000 words
Personal tonePerspective signals per 100 words
Section variationIntro-body-ending change
#1
30.3/100Low AI
0.59Low AI
0.0056Low AI
0.891Low AI
102.8/1kLow AI
9.8%Low AI
0.297Low AI
#2
30.9/100Low AI
0.66Low AI
0.0001Low AI
0.885Medium AI
95.5/1kLow AI
5.9%High AI
0.242Low AI
#3
35.3/100Medium AI
0.64Low AI
0.0284Low AI
0.887Low AI
84/1kLow AI
4.7%High AI
0.244Low AI
#4
36.9/100Medium AI
0.56Low AI
0.0004Low AI
0.889Low AI
78.4/1kLow AI
4.9%High AI
0.235Medium AI
#5
39.9/100Medium AI
0.65Low AI
0Low AI
0.889Low AI
111/1kLow AI
7.1%High AI
0.209High AI
#6
41.5/100Medium AI
0.49High AI
0.0051Low AI
0.874High AI
104.1/1kLow AI
5.5%High AI
0.215High AI
#7
42.3/100Medium AI
0.55Low AI
0.0087Low AI
0.838High AI
73.7/1kLow AI
3.9%High AI
0.243Low AI
#8
43.1/100Medium AI
0.5High AI
0.007Low AI
0.884Medium AI
88.3/1kLow AI
5.7%High AI
0.197High AI
#9
43.3/100Medium AI
0.6Low AI
0.0145Low AI
0.887Low AI
109.1/1kLow AI
7.6%High AI
0.213High AI
#10
44.3/100Medium AI
0.65Low AI
0Low AI
0.88High AI
115.3/1kMedium AI
6.9%High AI
0.23Medium AI
#11
48.5/100Medium AI
0.52High AI
0.0033Low AI
0.882High AI
118.9/1kMedium AI
7.3%High AI
0.199High AI
#12
51/100High AI
0.53Medium AI
0.0021Low AI
0.856High AI
115.8/1kMedium AI
7.9%High AI
0.225High AI
#13
52.4/100High AI
0.48High AI
0.9996High AI
0.894Low AI
90.8/1kLow AI
4.8%High AI
0.293Low AI
#14
53.3/100High AI
0.49High AI
0.0212Low AI
0.876High AI
112.1/1kLow AI
5.7%High AI
0.17High AI
#15
54.1/100High AI
0.57Low AI
0.3335High AI
0.873High AI
85.1/1kLow AI
6.7%High AI
0.23Medium AI
#16
54.6/100High AI
0.51High AI
0.024Low AI
0.871High AI
104.9/1kLow AI
5.3%High AI
0.154High AI
#17
55.9/100High AI
0.49High AI
0.0157Low AI
0.874High AI
115.3/1kMedium AI
7.1%High AI
0.214High AI
#18
56.1/100High AI
0.5High AI
0.0003Low AI
0.866High AI
111/1kLow AI
5.7%High AI
0.218High AI
#19
58/100High AI
0.56Low AI
0.036Low AI
0.873High AI
123.9/1kHigh AI
7%High AI
0.2High AI
#20
59.5/100High AI
0.5High AI
0.0295Low AI
0.875High AI
105.7/1kLow AI
5%High AI
0.157High AI
#21
60.5/100High AI
0.42High AI
0.1528High AI
0.886Medium AI
122.1/1kHigh AI
7%High AI
0.238Low AI
#22
60.9/100High AI
0.51High AI
0.629High AI
0.875High AI
108/1kLow AI
10%Low AI
0.207High AI
#23
62.1/100High AI
0.57Low AI
0.1023High AI
0.866High AI
111.3/1kLow AI
10.6%Low AI
0.195High AI
#24
71.8/100High AI
0.43High AI
0.2718High AI
0.877High AI
130.4/1kHigh AI
7.9%High AI
0.143High AI
#25
73.4/100High AI
0.29High AI
0.9944High AI
0.893Low AI
135.4/1kHigh AI
7.6%High AI
0.213High AI
Metric labels: Low AI through 25% evidence, Medium through 50%, High above 50%. Global score: Low through 35, Medium through 50, High above 50.
How to read it
Style evidence, not authorship proof
The AI Score is the equal-weight average of six interpretable writing signals. Use it to compare model tendencies under controlled tasks, not to determine who wrote an isolated text.
1
Start with the AI Score
Lower scores rank higher. The score equally averages all six calibrated signals for every completed writing task.
2
Compare the metric labels
Green is near the low-evidence end of a calibration. Yellow is intermediate; red has crossed the interval midpoint.
3
Open a model, then a task
Model analytics show averages. Task analytics reveal genre-specific scores, graphs, recommendations and highlighted output.
Benchmark methodology
Six writing signals used in every AI Score
Every model is evaluated with the same metrics and equal weights. The final AI Score is the average of the six calibrated evidence values across that model's completed tasks.
Exact score calculation: task AI Score = 100 x mean(the six available 0-1 evidence values). Model AI Score = mean(all completed task AI Scores). Individual raw values are averaged only for display; the model score is not recalculated from those displayed averages.
Sentence rhythm
Counts the words in every sentence, calculates their population standard deviation, then divides it by the mean sentence length. A larger coefficient means a wider mix of short and long sentences.
Formula: SD(sentence word counts) / mean sentence word count. Evidence anchors: 0.55 = 0; 0.50 = 1.
GPT-2 perplexity proximity
DistilGPT-2 compares the observed next-token log probability with the probability distribution it expected at every position. The standardized difference is converted through the normal cumulative distribution.
spaCy converts the text into part-of-speech sequences. The score combines normalized entropy for single tags, two-tag sequences, three-tag sequences, and unique sentence-level POS patterns.
Formula: 0.25 x unigram entropy + 0.35 x bigram entropy + 0.35 x trigram entropy + 0.05 x sentence-pattern diversity. Evidence anchors: 0.890 = 0; 0.875 = 1.
Prose heaviness
Counts nominalizations, reduced participial clauses, subject-that constructions, and phrasal coordination, then normalizes their total for text length.
Formula: dense constructions / total words x 1,000. Evidence anchors: 110/1k = 0; 130/1k = 1.
Personal tone
Counts personal pronouns, private-state verbs such as think or feel, hedges, and direct questions. It estimates visible perspective, uncertainty, and reader engagement.
Formula: (personal pronouns + private verbs + hedges + questions) / total words x 100. Evidence anchors: 9% = 0; 7.5% = 1.
Section variation
Splits the text into opening, middle, and ending thirds. Each third is profiled by sentence length, sentence rhythm, function-word rate, and modifier rate; normalized pairwise distances are averaged.
Formula: mean pairwise distance among the three section profiles. Evidence anchors: 0.25 = 0; 0.20 = 1.
Controlled comparison
Eight writing genres
Using different genres reduces the chance that one topic or writing format determines the leaderboard.
SEO article about coffee roasting
Academic essay on universal basic income
Research writing about AI detection
Personal writing about overcoming insecurity
Fiction about a Halloween party going wrong
Biography of Alan Turing
Copywriting email for a dating service
News report on New York's cost of living
Questions
About the LLM writing benchmark
Does a high score prove that a text was written by AI?
No. The score compares measurable stylistic tendencies. It is useful for controlled model comparison, but it is not proof of authorship for an isolated document.
Why is the human baseline included?
The human samples provide a practical reference under the same eight genres. They make the direction and scale of each model result easier to interpret.
How often does the leaderboard update?
The website reads the latest benchmark published from the Intellectualead app. Completed NLP backfills and new benchmark runs publish refreshed metrics automatically.