Sentence length diversity
Looks at whether a text naturally mixes short, medium and long sentences, including the center and tails of the distribution. Broader, less repetitive pacing is treated as more human-like.
Independent style analysis
Compare frontier language models writing quality based on stylistic metrics (grammatical, syntactic, rethorical, lexical, prosody...) and with human writing as a reference. Lower AI evidence means the model output is stylistically closer to the human baseline used in this benchmark. Since the leaderboard is based on a stylistic analysis and not blind human preferences or machine-learning discrimination like AI detectors, you can understand why a model sounds more human than another.
Scores update automatically when a new benchmark is published.
| Rank | ModelLatest published runs | AI ScoreAverage of nine signals | Sentence length diversityDistribution breadth, quantiles and tails | GPT-2 perplexity similarityModel-like token predictability | Noun contextualizationContext supplied around nouns | Heavy grammarLLM-favored constructions per 1,000 lexical tokens | Personal tonePerspective signals per 100 words | AI SlopRhetorical templates per 1,000 words | AI WordingOverused wording per 1,000 words | Punctuation richnessRare-sensitive variety across ten mark families | Section variationIntro-body-ending change |
|---|---|---|---|---|---|---|---|---|---|---|---|
| #1 | 16.2/100Low AICombined evidence is 35/100 or lower. | 0.744Low AISentence lengths are broadly distributed. | 0.0139Low AIToken choices are within the lowest evidence quarter. | 0.853Low AINouns are well contextualized with linking, helping, and qualifying language. | 13.3/1kLow AIFew of the directional dense-grammar signals are elevated. | 7.3%Low AIPersonal-tone density is near the human anchor. | 6.5/1kLow AI6.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 5.4/1kLow AI6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.766Low AIPunctuation is richly distributed across common and less-common mark types. | 0.31Low AIThe document changes shape naturally. | |
| #2 | 27.3/100Low AICombined evidence is 35/100 or lower. | 0.711Medium AISentence lengths use several bands. | 0Low AIToken choices are within the lowest evidence quarter. | 0.759Low AINouns are well contextualized with linking, helping, and qualifying language. | 15.1/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.9%High AIVery little visible perspective or uncertainty. | 11.1/1kMedium AI12.5 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 4.5/1kLow AI5 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.726Low AIPunctuation is richly distributed across common and less-common mark types. | 0.221High AIVery little style change across the document. | |
| #3 | 28.9/100Low AICombined evidence is 35/100 or lower. | 0.708Medium AISentence lengths use several bands. | 0Low AIToken choices are within the lowest evidence quarter. | 0.8Low AINouns are well contextualized with linking, helping, and qualifying language. | 18.7/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.9%High AIVery little visible perspective or uncertainty. | 11.5/1kMedium AI12.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 3.6/1kLow AI3.8 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.668Medium AISome less-common punctuation types are represented. | 0.207High AIVery little style change across the document. | |
| #4 | 29.3/100Low AICombined evidence is 35/100 or lower. | 0.668High AISentence lengths are concentrated in a few bands. | 0.0013Low AIToken choices are within the lowest evidence quarter. | 0.808Low AINouns are well contextualized with linking, helping, and qualifying language. | 15.1/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.8%High AIVery little visible perspective or uncertainty. | 5.5/1kLow AI6.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 5.1/1kLow AI5.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.56High AIPunctuation is concentrated in relatively few mark types. | 0.231Medium AISome section-to-section change. | |
| #5 | 30.2/100Low AICombined evidence is 35/100 or lower. | 0.672High AISentence lengths are concentrated in a few bands. | 0.001Low AIToken choices are within the lowest evidence quarter. | 0.608High AINouns receive little relational or qualifying context. | 14.1/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.8%High AIVery little visible perspective or uncertainty. | 10.8/1kLow AI16.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 7.5/1kLow AI11 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.758Low AIPunctuation is richly distributed across common and less-common mark types. | 0.303Low AIThe document changes shape naturally. | |
| #6 | 33.2/100Low AICombined evidence is 35/100 or lower. | 0.755Low AISentence lengths are broadly distributed. | 0Low AIToken choices are within the lowest evidence quarter. | 0.779Low AINouns are well contextualized with linking, helping, and qualifying language. | 18.5/1kLow AIFew of the directional dense-grammar signals are elevated. | 5%High AIVery little visible perspective or uncertainty. | 13.2/1kMedium AI14.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 4.9/1kLow AI5.5 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.596High AIPunctuation is concentrated in relatively few mark types. | 0.209High AIVery little style change across the document. | |
| #7 | 34.6/100Low AICombined evidence is 35/100 or lower. | 0.701Medium AISentence lengths use several bands. | 0.0001Low AIToken choices are within the lowest evidence quarter. | 0.741Low AINouns are well contextualized with linking, helping, and qualifying language. | 17.8/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.1%High AIVery little visible perspective or uncertainty. | 13.9/1kMedium AI14.4 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 5.1/1kLow AI5.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.609High AIPunctuation is concentrated in relatively few mark types. | 0.249Medium AISome section-to-section change. | |
| #8 | 35.7/100Medium AICombined evidence is above 35 and no higher than 50. | 0.604High AISentence lengths are concentrated in a few bands. | 0.0004Low AIToken choices are within the lowest evidence quarter. | 0.686Medium AINoun contextualization is moderate. | 19.8/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.6%High AIVery little visible perspective or uncertainty. | 7.1/1kLow AI7.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 7.8/1kLow AI8.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.646Medium AISome less-common punctuation types are represented. | 0.218High AIVery little style change across the document. | |
| #9 | 36/100Medium AICombined evidence is above 35 and no higher than 50. | 0.672High AISentence lengths are concentrated in a few bands. | 0.0145Low AIToken choices are within the lowest evidence quarter. | 0.784Low AINouns are well contextualized with linking, helping, and qualifying language. | 19.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.2%Medium AISome personal or subjective language. | 12.9/1kMedium AI13.3 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 8.5/1kLow AI8.9 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.58High AIPunctuation is concentrated in relatively few mark types. | 0.213High AIVery little style change across the document. | |
| #10 | 36.1/100Medium AICombined evidence is above 35 and no higher than 50. | 0.616High AISentence lengths are concentrated in a few bands. | 0.0157Low AIToken choices are within the lowest evidence quarter. | 0.747Low AINouns are well contextualized with linking, helping, and qualifying language. | 17.7/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.3%Medium AISome personal or subjective language. | 9.5/1kLow AI13 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 8.2/1kLow AI12.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.605High AIPunctuation is concentrated in relatively few mark types. | 0.217High AIVery little style change across the document. | |
| #11 | 36.4/100Medium AICombined evidence is above 35 and no higher than 50. | 0.6High AISentence lengths are concentrated in a few bands. | 0.0282Low AIToken choices are within the lowest evidence quarter. | 0.845Low AINouns are well contextualized with linking, helping, and qualifying language. | 14.6/1kLow AIFew of the directional dense-grammar signals are elevated. | 7.2%Low AIPersonal-tone density is near the human anchor. | 8.7/1kLow AI8.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 3.8/1kLow AI3.9 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.603High AIPunctuation is concentrated in relatively few mark types. | 0.179High AIVery little style change across the document. | |
| #12 | 36.9/100Medium AICombined evidence is above 35 and no higher than 50. | 0.727Low AISentence lengths are broadly distributed. | 0.0897Medium AIToken choices are between one quarter and one half of the interval. | 0.746Low AINouns are well contextualized with linking, helping, and qualifying language. | 22/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.7%High AIVery little visible perspective or uncertainty. | 12.2/1kMedium AI16.4 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 8.4/1kLow AI11.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.696Low AIPunctuation is richly distributed across common and less-common mark types. | 0.182High AIVery little style change across the document. | |
| #13 | 37/100Medium AICombined evidence is above 35 and no higher than 50. | 0.673High AISentence lengths are concentrated in a few bands. | 0.0284Low AIToken choices are within the lowest evidence quarter. | 0.691Medium AINoun contextualization is moderate. | 21.5/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.3%High AIVery little visible perspective or uncertainty. | 14.3/1kHigh AI14.9 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 10.6/1kMedium AI11.4 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.684Medium AISome less-common punctuation types are represented. | 0.244Medium AISome section-to-section change. | |
| #14 | 37.1/100Medium AICombined evidence is above 35 and no higher than 50. | 0.66High AISentence lengths are concentrated in a few bands. | 0.036Low AIToken choices are within the lowest evidence quarter. | 0.747Low AINouns are well contextualized with linking, helping, and qualifying language. | 21.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.8%High AIVery little visible perspective or uncertainty. | 8.9/1kLow AI13.5 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 10.9/1kMedium AI16.9 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.631Medium AISome less-common punctuation types are represented. | 0.203High AIVery little style change across the document. | |
| #15 | 37.5/100Medium AICombined evidence is above 35 and no higher than 50. | 0.634High AISentence lengths are concentrated in a few bands. | 0Low AIToken choices are within the lowest evidence quarter. | 0.677Medium AINoun contextualization is moderate. | 14.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 6%Medium AISome personal or subjective language. | 13.6/1kMedium AI18.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 8.7/1kLow AI11.5 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.718Low AIPunctuation is richly distributed across common and less-common mark types. | 0.202High AIVery little style change across the document. | |
| #16 | 37.6/100Medium AICombined evidence is above 35 and no higher than 50. | 0.667High AISentence lengths are concentrated in a few bands. | 0.0033Low AIToken choices are within the lowest evidence quarter. | 0.821Low AINouns are well contextualized with linking, helping, and qualifying language. | 22.6/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.5%High AIVery little visible perspective or uncertainty. | 11.3/1kMedium AI12 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 9.7/1kMedium AI10.4 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.56High AIPunctuation is concentrated in relatively few mark types. | 0.199High AIVery little style change across the document. | |
| #17 | 37.7/100Medium AICombined evidence is above 35 and no higher than 50. | 0.725Low AISentence lengths are broadly distributed. | 0.0016Low AIToken choices are within the lowest evidence quarter. | 0.553High AINouns receive little relational or qualifying context. | 23.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.3%High AIVery little visible perspective or uncertainty. | 9.6/1kLow AI11.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 14.5/1kHigh AI18.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.74Low AIPunctuation is richly distributed across common and less-common mark types. | 0.252Medium AISome section-to-section change. | |
| #18 | 37.9/100Medium AICombined evidence is above 35 and no higher than 50. | 0.707Medium AISentence lengths use several bands. | 0.0004Low AIToken choices are within the lowest evidence quarter. | 0.836Low AINouns are well contextualized with linking, helping, and qualifying language. | 22/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.8%High AIVery little visible perspective or uncertainty. | 15/1kHigh AI15.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 8.2/1kLow AI8.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.531High AIPunctuation is concentrated in relatively few mark types. | 0.235Medium AISome section-to-section change. | |
| #19 | 38.4/100Medium AICombined evidence is above 35 and no higher than 50. | 0.683High AISentence lengths are concentrated in a few bands. | 0Low AIToken choices are within the lowest evidence quarter. | 0.71Medium AINoun contextualization is moderate. | 19.6/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.4%High AIVery little visible perspective or uncertainty. | 15.5/1kHigh AI16.6 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 8/1kLow AI8.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.634Medium AISome less-common punctuation types are represented. | 0.23Medium AISome section-to-section change. | |
| #20 | 38.7/100Medium AICombined evidence is above 35 and no higher than 50. | 0.709Medium AISentence lengths use several bands. | 0.0087Low AIToken choices are within the lowest evidence quarter. | 0.49High AINouns receive little relational or qualifying context. | 22.6/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.5%High AIVery little visible perspective or uncertainty. | 9.8/1kLow AI11.4 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 18/1kHigh AI20.9 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.733Low AIPunctuation is richly distributed across common and less-common mark types. | 0.247Medium AISome section-to-section change. | |
| #21 | 38.8/100Medium AICombined evidence is above 35 and no higher than 50. | 0.699Medium AISentence lengths use several bands. | 0.1031High AIToken choices are beyond the interval midpoint. | 0.737Low AINouns are well contextualized with linking, helping, and qualifying language. | 27.9/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.6%High AIVery little visible perspective or uncertainty. | 9.3/1kLow AI10.3 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 20.9/1kHigh AI23.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.692Low AIPunctuation is richly distributed across common and less-common mark types. | 0.181High AIVery little style change across the document. | |
| #22 | 38.9/100Medium AICombined evidence is above 35 and no higher than 50. | 0.743Low AISentence lengths are broadly distributed. | 0.0104Low AIToken choices are within the lowest evidence quarter. | 0.534High AINouns receive little relational or qualifying context. | 24.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 4%High AIVery little visible perspective or uncertainty. | 9.7/1kLow AI11.4 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 14.2/1kHigh AI17.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.663Medium AISome less-common punctuation types are represented. | 0.272Low AIThe document changes shape naturally. | |
| #23 | 39.1/100Medium AICombined evidence is above 35 and no higher than 50. | 0.635High AISentence lengths are concentrated in a few bands. | 0Low AIToken choices are within the lowest evidence quarter. | 0.787Low AINouns are well contextualized with linking, helping, and qualifying language. | 17.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.9%High AIVery little visible perspective or uncertainty. | 10.1/1kLow AI13.6 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 3.7/1kLow AI5.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.528High AIPunctuation is concentrated in relatively few mark types. | 0.168High AIVery little style change across the document. | |
| #24 | 39.4/100Medium AICombined evidence is above 35 and no higher than 50. | 0.695High AISentence lengths are concentrated in a few bands. | 0.0021Low AIToken choices are within the lowest evidence quarter. | 0.567High AINouns receive little relational or qualifying context. | 20.3/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.7%High AIVery little visible perspective or uncertainty. | 13.3/1kMedium AI18.5 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 10.7/1kMedium AI15.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.682Medium AISome less-common punctuation types are represented. | 0.23Medium AISome section-to-section change. | |
| #25 | 39.4/100Medium AICombined evidence is above 35 and no higher than 50. | 0.714Medium AISentence lengths use several bands. | 0.0011Low AIToken choices are within the lowest evidence quarter. | 0.651Medium AINoun contextualization is moderate. | 22.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.3%High AIVery little visible perspective or uncertainty. | 17.6/1kHigh AI24.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 10.8/1kMedium AI17 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.731Low AIPunctuation is richly distributed across common and less-common mark types. | 0.214High AIVery little style change across the document. | |
| #26 | 41.1/100Medium AICombined evidence is above 35 and no higher than 50. | 0.672High AISentence lengths are concentrated in a few bands. | 0.9996High AIToken choices are beyond the interval midpoint. | 0.737Low AINouns are well contextualized with linking, helping, and qualifying language. | 18.5/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.7%High AIVery little visible perspective or uncertainty. | 5.6/1kLow AI4.3 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 15.6/1kHigh AI11.4 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.688Medium AISome less-common punctuation types are represented. | 0.293Low AIThe document changes shape naturally. | |
| #27 | 41.3/100Medium AICombined evidence is above 35 and no higher than 50. | 0.676High AISentence lengths are concentrated in a few bands. | 0.0002Low AIToken choices are within the lowest evidence quarter. | 0.551High AINouns receive little relational or qualifying context. | 26.8/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.7%High AIVery little visible perspective or uncertainty. | 8.8/1kLow AI8.9 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 27.9/1kHigh AI33.4 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.726Low AIPunctuation is richly distributed across common and less-common mark types. | 0.259Medium AISome section-to-section change. | |
| #28 | 41.6/100Medium AICombined evidence is above 35 and no higher than 50. | 0.626High AISentence lengths are concentrated in a few bands. | 0.0295Low AIToken choices are within the lowest evidence quarter. | 0.766Low AINouns are well contextualized with linking, helping, and qualifying language. | 20.5/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.1%Medium AISome personal or subjective language. | 9.5/1kLow AI12 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 7.7/1kLow AI10.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.566High AIPunctuation is concentrated in relatively few mark types. | 0.159High AIVery little style change across the document. | |
| #29 | 42.3/100Medium AICombined evidence is above 35 and no higher than 50. | 0.636High AISentence lengths are concentrated in a few bands. | 0.007Low AIToken choices are within the lowest evidence quarter. | 0.656Medium AINoun contextualization is moderate. | 24.8/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.4%Medium AISome personal or subjective language. | 13.7/1kMedium AI13.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 15.7/1kHigh AI14.9 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.651Medium AISome less-common punctuation types are represented. | 0.197High AIVery little style change across the document. | |
| #30 | 42.5/100Medium AICombined evidence is above 35 and no higher than 50. | 0.579High AISentence lengths are concentrated in a few bands. | 0.0051Low AIToken choices are within the lowest evidence quarter. | 0.638High AINouns receive little relational or qualifying context. | 21.1/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.7%Low AIPersonal-tone density is near the human anchor. | 9.6/1kLow AI10.9 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 6.9/1kLow AI7.9 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.625Medium AISome less-common punctuation types are represented. | 0.213High AIVery little style change across the document. | |
| #31 | 43.2/100Medium AICombined evidence is above 35 and no higher than 50. | 0.557High AISentence lengths are concentrated in a few bands. | 0Low AIToken choices are within the lowest evidence quarter. | 0.616High AINouns receive little relational or qualifying context. | 22.3/1kLow AIFew of the directional dense-grammar signals are elevated. | 7%Low AIPersonal-tone density is near the human anchor. | 8.4/1kLow AI8.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 7/1kLow AI7.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.589High AIPunctuation is concentrated in relatively few mark types. | 0.181High AIVery little style change across the document. | |
| #32 | 43.7/100Medium AICombined evidence is above 35 and no higher than 50. | 0.615High AISentence lengths are concentrated in a few bands. | 0.0051Low AIToken choices are within the lowest evidence quarter. | 0.657Medium AINoun contextualization is moderate. | 30.3/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.4%Medium AISome personal or subjective language. | 11.1/1kMedium AI11.5 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 13.9/1kHigh AI14.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.636Medium AISome less-common punctuation types are represented. | 0.215High AIVery little style change across the document. | |
| #33 | 43.8/100Medium AICombined evidence is above 35 and no higher than 50. | 0.677High AISentence lengths are concentrated in a few bands. | 0.629High AIToken choices are beyond the interval midpoint. | 0.791Low AINouns are well contextualized with linking, helping, and qualifying language. | 23/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.2%High AIVery little visible perspective or uncertainty. | 10.5/1kLow AI11.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 17.5/1kHigh AI20.8 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.694Low AIPunctuation is richly distributed across common and less-common mark types. | 0.209High AIVery little style change across the document. | |
| #34 | 43.9/100Medium AICombined evidence is above 35 and no higher than 50. | 0.669High AISentence lengths are concentrated in a few bands. | 0.0116Low AIToken choices are within the lowest evidence quarter. | 0.613High AINouns receive little relational or qualifying context. | 21.7/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.7%High AIVery little visible perspective or uncertainty. | 10.2/1kLow AI8.9 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 10.5/1kMedium AI8.9 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.563High AIPunctuation is concentrated in relatively few mark types. | 0.194High AIVery little style change across the document. | |
| #35 | 44/100Medium AICombined evidence is above 35 and no higher than 50. | 0.577High AISentence lengths are concentrated in a few bands. | 0.0001Low AIToken choices are within the lowest evidence quarter. | 0.608High AINouns receive little relational or qualifying context. | 24.5/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.7%Low AIPersonal-tone density is near the human anchor. | 5.8/1kLow AI6 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 5.5/1kLow AI5.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.552High AIPunctuation is concentrated in relatively few mark types. | 0.194High AIVery little style change across the document. | |
| #36 | 44.2/100Medium AICombined evidence is above 35 and no higher than 50. | 0.662High AISentence lengths are concentrated in a few bands. | 0.02Low AIToken choices are within the lowest evidence quarter. | 0.719Medium AINoun contextualization is moderate. | 24.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.1%High AIVery little visible perspective or uncertainty. | 14.1/1kHigh AI15.8 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 10.8/1kMedium AI11.5 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.601High AIPunctuation is concentrated in relatively few mark types. | 0.175High AIVery little style change across the document. | |
| #37 | 44.4/100Medium AICombined evidence is above 35 and no higher than 50. | 0.658High AISentence lengths are concentrated in a few bands. | 0.0028Low AIToken choices are within the lowest evidence quarter. | 0.605High AINouns receive little relational or qualifying context. | 25.5/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.8%High AIVery little visible perspective or uncertainty. | 12.7/1kMedium AI14.4 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 8.8/1kLow AI10 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.659Medium AISome less-common punctuation types are represented. | 0.189High AIVery little style change across the document. | |
| #38 | 44.5/100Medium AICombined evidence is above 35 and no higher than 50. | 0.652High AISentence lengths are concentrated in a few bands. | 0.349High AIToken choices are beyond the interval midpoint. | 0.781Low AINouns are well contextualized with linking, helping, and qualifying language. | 23/1kLow AIFew of the directional dense-grammar signals are elevated. | 6%Medium AISome personal or subjective language. | 10.2/1kLow AI10.9 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 20/1kHigh AI21 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.647Medium AISome less-common punctuation types are represented. | 0.182High AIVery little style change across the document. | |
| #39 | 45/100Medium AICombined evidence is above 35 and no higher than 50. | 0.685High AISentence lengths are concentrated in a few bands. | 0.3445High AIToken choices are beyond the interval midpoint. | 0.777Low AINouns are well contextualized with linking, helping, and qualifying language. | 25.1/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.3%High AIVery little visible perspective or uncertainty. | 11.9/1kMedium AI13.5 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 21.1/1kHigh AI24.3 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.637Medium AISome less-common punctuation types are represented. | 0.217High AIVery little style change across the document. | |
| #40 | 45.1/100Medium AICombined evidence is above 35 and no higher than 50. | 0.654High AISentence lengths are concentrated in a few bands. | 0.2912High AIToken choices are beyond the interval midpoint. | 0.689Medium AINoun contextualization is moderate. | 24.6/1kLow AIFew of the directional dense-grammar signals are elevated. | 5%High AIVery little visible perspective or uncertainty. | 12.1/1kMedium AI11.9 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 20.9/1kHigh AI20 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.702Low AIPunctuation is richly distributed across common and less-common mark types. | 0.216High AIVery little style change across the document. | |
| #41 | 46.1/100Medium AICombined evidence is above 35 and no higher than 50. | 0.722Medium AISentence lengths use several bands. | 0.6349High AIToken choices are beyond the interval midpoint. | 0.772Low AINouns are well contextualized with linking, helping, and qualifying language. | 23.2/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.2%High AIVery little visible perspective or uncertainty. | 15.6/1kHigh AI15.3 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 15.3/1kHigh AI14.5 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.703Low AIPunctuation is richly distributed across common and less-common mark types. | 0.187High AIVery little style change across the document. | |
| #42 | 46.4/100Medium AICombined evidence is above 35 and no higher than 50. | 0.669High AISentence lengths are concentrated in a few bands. | 0Low AIToken choices are within the lowest evidence quarter. | 0.483High AINouns receive little relational or qualifying context. | 25.5/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.5%High AIVery little visible perspective or uncertainty. | 10/1kLow AI11.3 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 19.7/1kHigh AI21.8 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.654Medium AISome less-common punctuation types are represented. | 0.226High AIVery little style change across the document. | |
| #43 | 46.7/100Medium AICombined evidence is above 35 and no higher than 50. | 0.661High AISentence lengths are concentrated in a few bands. | 0.1528High AIToken choices are beyond the interval midpoint. | 0.723Medium AINoun contextualization is moderate. | 23.8/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.2%High AIVery little visible perspective or uncertainty. | 10.6/1kLow AI10.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 22.9/1kHigh AI24.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.609High AIPunctuation is concentrated in relatively few mark types. | 0.236Medium AISome section-to-section change. | |
| #44 | 47.4/100Medium AICombined evidence is above 35 and no higher than 50. | 0.647High AISentence lengths are concentrated in a few bands. | 0.0007Low AIToken choices are within the lowest evidence quarter. | 0.614High AINouns receive little relational or qualifying context. | 24.7/1kLow AIFew of the directional dense-grammar signals are elevated. | 4.8%High AIVery little visible perspective or uncertainty. | 20.8/1kHigh AI19.6 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 23.5/1kHigh AI22.8 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.757Low AIPunctuation is richly distributed across common and less-common mark types. | 0.21High AIVery little style change across the document. | |
| #45 | 47.6/100Medium AICombined evidence is above 35 and no higher than 50. | 0.706Medium AISentence lengths use several bands. | 0.0128Low AIToken choices are within the lowest evidence quarter. | 0.597High AINouns receive little relational or qualifying context. | 33.7/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.4%High AIVery little visible perspective or uncertainty. | 15/1kHigh AI15.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 18.3/1kHigh AI17.5 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.609High AIPunctuation is concentrated in relatively few mark types. | 0.191High AIVery little style change across the document. | |
| #46 | 48.8/100Medium AICombined evidence is above 35 and no higher than 50. | 0.679High AISentence lengths are concentrated in a few bands. | 0.2718High AIToken choices are beyond the interval midpoint. | 0.735Low AINouns are well contextualized with linking, helping, and qualifying language. | 22.6/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.6%High AIVery little visible perspective or uncertainty. | 12.1/1kMedium AI15 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 26.8/1kHigh AI34.1 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.726Low AIPunctuation is richly distributed across common and less-common mark types. | 0.152High AIVery little style change across the document. | |
| #47 | 49.1/100Medium AICombined evidence is above 35 and no higher than 50. | 0.569High AISentence lengths are concentrated in a few bands. | 0.024Low AIToken choices are within the lowest evidence quarter. | 0.629High AINouns receive little relational or qualifying context. | 20.4/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.6%Medium AISome personal or subjective language. | 9.9/1kLow AI10.3 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 9.7/1kMedium AI9.9 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.521High AIPunctuation is concentrated in relatively few mark types. | 0.152High AIVery little style change across the document. | |
| #48 | 51/100High AICombined evidence is above 50. | 0.642High AISentence lengths are concentrated in a few bands. | 0.0001Low AIToken choices are within the lowest evidence quarter. | 0.592High AINouns receive little relational or qualifying context. | 39/1kMedium AISome dense constructions are elevated. | 6%Medium AISome personal or subjective language. | 15.5/1kHigh AI15.4 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 14.2/1kHigh AI13.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.619Medium AISome less-common punctuation types are represented. | 0.22High AIVery little style change across the document. | |
| #49 | 51.9/100High AICombined evidence is above 50. | 0.663High AISentence lengths are concentrated in a few bands. | 0.9944High AIToken choices are beyond the interval midpoint. | 0.79Low AINouns are well contextualized with linking, helping, and qualifying language. | 18.5/1kLow AIFew of the directional dense-grammar signals are elevated. | 6.2%Medium AISome personal or subjective language. | 5.2/1kLow AI3.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 16.4/1kHigh AI9.6 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.441High AIPunctuation is concentrated in relatively few mark types. | 0.213High AIVery little style change across the document. | |
| #50 | 55.8/100High AICombined evidence is above 50. | 0.61High AISentence lengths are concentrated in a few bands. | 0.4131High AIToken choices are beyond the interval midpoint. | 0.588High AINouns receive little relational or qualifying context. | 21.8/1kLow AIFew of the directional dense-grammar signals are elevated. | 5.3%High AIVery little visible perspective or uncertainty. | 13.8/1kMedium AI11.1 average matches. Evidence is calibrated from 8 to 20 matches per 1,000 words. | 26.4/1kHigh AI23.1 average list matches. Evidence is calibrated from 5 to 22 matches per 1,000 words. | 0.645Medium AISome less-common punctuation types are represented. | 0.266Low AIThe document changes shape naturally. |
Benchmark methodology
Every model is evaluated with the same nine stylistic signals and the same set of writing genres. The score summarizes those signals across each model's completed tasks.
Looks at whether a text naturally mixes short, medium and long sentences, including the center and tails of the distribution. Broader, less repetitive pacing is treated as more human-like.
Compares token predictability with patterns associated with language-model writing. It is one probabilistic signal, not a standalone detector.
Divides adpositions, auxiliaries and adverbs by nouns. A higher value means noun-led ideas receive more relational, tense or qualifying context. The calibration gives no AI evidence at 0.80 and full evidence at 0.50.
Tracks dense constructions including participial clauses, that-subject clauses, and 2-4 word adjective/noun sequences ending in an abstract -ty/-ity, -tion/-sion, -ness, or -ment noun.
Looks for four stance families: self-mentions including I, we, you, and your; hedges; boosters; and attitude markers, plus direct questions. Modal auxiliaries and multiword phrases such as “in practice” or “in fact” are classified inside those families and are not counted twice. A value of 7.4% gives zero AI evidence and 4.6% gives full AI evidence.
Tracks rhetorical templates such as em dashes, three-part sequences, formulaic contrasts, while-led concessions, generic setups, unsupported attribution and recap endings. Occasional use is normal; repetition is more informative.
Tracks words and short phrases that appear unusually often in LLM output. The metric is most useful when several generic matches accumulate in one text.
Checks whether the opening, body and ending perform differently or repeat the same stylistic shape. More purposeful change is treated as more human-like.
Uses rare-sensitive Hill diversity at q = 0.25 across ten punctuation families. It rewards meaningful representation of less-common marks more strongly than ordinary entropy and contributes equally to the AI Score.
Controlled comparison
Using different genres reduces the chance that one topic or writing format determines the leaderboard.
Questions
It summarizes eight interpretable stylistic signals across the same controlled writing tasks. A lower result means the model's outputs are stylistically closer to the human reference used here; it does not mean every individual passage will sound human.
No. The score compares measurable stylistic tendencies. It is useful for controlled model comparison, but it is not proof of authorship for an isolated document.
The human samples provide a consistent reference under the same eight genres and tasks. They make the direction and scale of each model result easier to interpret without claiming that one human style is universally ideal.
Preference leaderboards report which answer readers or judges favor. This benchmark measures explainable properties of the writing itself, including grammar, syntax, rhetorical habits, wording, sentence pacing and section structure.
AI detectors usually estimate whether an isolated text belongs to an AI or human class. This benchmark compares models under controlled prompts and exposes the stylistic signals behind the ranking. It should not be used to accuse an author or verify authorship.
Writing style changes with purpose. Testing SEO, academic, research, personal, fiction, biography, email and news writing reduces the chance that one topic or format decides the ranking.
A ranking may change after new task outputs are published, a provider updates a model, missing analyses are completed, or the benchmark's linguistic analysis is improved. The table always uses the latest published runs.
Yes. Their overall scores can be close while their metric profiles differ substantially. Open the model details to compare which model relies more on dense grammar, formulaic rhetoric, predictable wording or uniform section structure.
The website reads the latest benchmark published from the Intellectualead app. Completed NLP backfills and new benchmark runs publish refreshed metrics automatically.