IntellectuaLead
English

Best LLMs for Writing in 2026 based on Leaderboards

You want to know what are the best LLM for creative, fictional, non-fictional writing ? Here's your guide.

Buchert Jean-marc

Buchert Jean-marc

August 6, 2026 • 18 min read
Free Prompts

Free Prompts and Ebook to Humanize Your Text

Download Now

AI can fill your blank page, but better pick the right large-language model (LLM).

Here, we give you the best LLM for writers based on a stylistic leaderboard and the Chatbot Arena leaderbaoard.

Muse AI by Sudowrite creative writing illustration

✨ Meet Muse — The LLM Built for Storytellers

If you’ve ever wished your LLM could write with real emotion, rhythm, and voiceMuse by Sudowrite was created exactly for that. It’s the only AI trained purely on outstanding fiction and creative prose, so your scenes, characters, and dialogue sound vivid and human.

Tested & Recommended by us!

🚀 Try Muse For Free Here

*No setup or credit card needed – just open Sudowrite and try Muse.*

How We Ranked the LLM Models

To find the best LLM models for writing, we used two leaderboards the Chatbot Arena, and our own stylistic leaderboard.

Chatbot Arena is a crowdsourced leaderboard run by LMSYS. It compares large language models based on blind user preferences across many tasks—including creative writing.

We use both the Creative Writing leaderboard and the broader text ranking. The creative board shows which outputs readers prefer as writing.

Then, there’s our stylistic benchmark that you can check out here.

In this leaderboard, each LLM model output are assessed through 8 writing tasks and 6 lexical, grammatical and stylistic metrics drawn from the stylometric academic litterature :

  • Sentence rhythm: assesses the standard deviation of sentence length divided by mean sentence length.
  • GPT-2 Perplexity Proximity: compares the text with DistilGPT-2’s expected token choices and probability profile.
  • Syntax diversity: analyzes the usage and distribution of grammatical elements (nouns, verbs, adverbs, adjectives, determiners etc…) and the variation of their combination.
  • Prose heaviness: counts nominalizations, reduced participial clauses, subject-that constructions, and phrasal coordination per 1,000 words.
  • Personal tone: counts personal pronouns, private-state verbs, hedges, and direct questions as a percentage of words.
  • Section variation: compares opening, middle, and ending thirds using sentence length, rhythm, function-word rate, and modifier rate.

I found that these metrics discriminate the best between human and LLM writing. And so are a perfect fit to assess a LLM model’s prose “humaness”.

1. Claude (Fable 5, Opus 5, Opus 4,8, Opus 4,7, Opus 4,6)

If I had to put one LLM family at the top for writing right now, I would still pick Claude.

The reason is simple. Claude still has the most life in the sentence. It is better than most models at keeping a voice stable, writing dialogue that sounds intentional, and carrying emotional tone across a full piece instead of just a paragraph.

Claude is not always the most rigid or procedural writer. For strict how-to content, technical walkthroughs, or very factual SEO sections, other models can still feel more exact. But that is also the tradeoff that makes Claude so useful. It usually gives you stronger rhythm, more natural phrasing, and less of that flat “AI explainer” feel.

The stylistic benchmark supports keeping Claude first. Claude Opus 4.6 scored 35.3 across all eight writing tests, and four Claude variants appeared within the benchmark’s first seven LLM positions.

In the current Chatbot Arena Creative writing snapshot, Claude models often rank the highest like Claude Fable 5, Opus 5 and Opus 4,7.

Technical specs

  • Context window (input): 1,000,000 tokens. Output window (single response): up to 128,000 tokens. Anthropic’s Claude 4.6 API docs list Opus 4.6 with a 1M-token context window, extended thinking, and a doubled 128K output cap.
  • Availability / hosting: Anthropic API, Google Vertex AI, Amazon Bedrock. Anthropic, Google Cloud, and AWS all list Opus 4.6 as available on their platforms.
  • Private/local: No local or offline self-hosting. Claude is delivered as a managed model through Anthropic or partner clouds rather than as open weights.

Pricing

Claude plans (Claude.ai): Free, Pro at $20/month or $17/month billed annually, and Max starting at $100/month. Anthropic’s pricing page says Pro includes more usage, access to more Claude models, and features like Research, Claude Code, and projects, while Max expands usage further.

API pricing (per million tokens):
Sonnet 4.6: $3 input / $15 output.
Opus 4.6: $5 input / $25 output.
Haiku 4.5: $1 input / $5 output. Anthropic also lists prompt caching discounts and batch-processing discounts, which can materially lower production costs.

2. Kimi (K3)

Kimi K3 and the Kimi model family are the biggest surprise of the last years in terms of writing quality.

Kimi writing has usually natural sentence rhythm, relatively light prose, and reflected writing structures. Kimi K3 looks unusually natural and less templated than most frontier models.

I still keep Kimi below Claude overall. Its personal-tone signal remains weaker, the Claude family has shown way more consistency over the years, and the model only ranks 24th on the Creative Writing category in Chatbot Arena. But still it’s very promising.

Our benchmark reference: Kimi K3 scored 30.9/100 across 7 tasks.

3. Meta (Muse 1,2, 1,1, Llama 3.1)

Meta Muse models have one of the most unusual writing profiles in the benchmark. Its grammar-pattern diversity and section-level variation look close to the human baseline, while its sentence rhythm and token predictability look much more model-like.

Meta Muse 1,2 scored 52.4 across all eight tasks.

The current Chatbot Arena rank for Muse Spark is 12th.

Technical specs

  • Model tested: Llama 3.1 70B Instruct (meta-llama/llama-3.1-70b-instruct).
  • Benchmark coverage: all eight writing tasks.
  • Deployment note: the benchmark identifies the model but does not supply context-window, quantization, or hosting details. Use the deployment documentation for the exact build you plan to run.

4. Gemini (3.6 Flash / 3.1 Pro)

Gemini is a complete LLM family for serious writing work. Its core strength is a mix of reasoning, long-context control, factual consistency, and a more procedural approach than Claude.

That makes it especially strong for SEO writing, essay writing, technical explainers, and research-heavy content where you need detailed structure without losing the thread. It is not always the most lyrical model for fiction, but it is one of the easiest models to trust with a complex brief.

In the stylistic benchmark, the latest Gemini 3.6 Flash scored 542.3 across all eight tests, clearly ahead of Gemini 3.1 Pro at 53.3 and Gemini 2.5 Pro at 60.9.

Gemini 3.1 Pro is also hitting #5 on Chatbot Arena creative writing leaderboard while Gemini-3.6-flash is #12 which is remarkable for a model of this size.

Technical specs

  • Context window (input): 1,048,576 tokens. Output window: 65,536 tokens. Model code: gemini-3.1-pro-preview.
  • Hosting options: Vertex AI, Google AI Studio / Gemini API, Gemini app, Gemini Enterprise, NotebookLM, Gemini CLI, and Google’s agentic development tools.
  • Native multimodal inputs: text, image, video, audio, and PDF. Output: text. That is useful when you want to pull structure from PDFs, slides, screenshots, reports, or long research files before drafting.
  • Capabilities: thinking, function calling, code execution, structured outputs, search grounding, URL context, batch API, and caching are all supported.

Pricing

Gemini API / AI Studio: Token-based pricing. For prompts up to 200K tokens, Google lists $1.25 per 1M input tokens and $10 per 1M output tokens. For prompts above 200K tokens, pricing rises to $2.50 input and $15 output per 1M tokens.

Context caching: $0.125 per 1M tokens for prompts up to 200K, and $0.25 per 1M above 200K, plus storage pricing. That matters if you reuse long style guides, product docs, or large research blocks across many drafts.

Grounding with Google Search: 1,500 grounded prompts per day free, then $35 per 1,000 grounded prompts.

Bundled subscriptions: Google AI Pro and Ultra plans include higher access to Gemini 3.1 Pro inside the Gemini app, alongside Deep Research and Workspace integrations. That is the easier route if your team mainly writes inside Google’s apps rather than through the API.

Vertex AI: Enterprise route for governance, quotas, and org-level controls. Pricing follows Google Cloud’s model-based billing. Use this lane when you need observability, security controls, or a more formal deployment path.

5. ChatGPT (GPT-5,6 Sol, 5,5)

Let’s be honest, the ChatGPT family are definitely not the best LLM models for writers. It has never been the safest choices for nonfiction, essay writing, SEO, and production content. It always feel engineered, with predictable transitions and less emotional movement across a full piece.

But for non-fiction and factual writing, it follows instructions well, handles structured formats cleanly, and is usually easier to steer than more expressive models.

The stylistic benchmark shows a clear improvement over older OpenAI models. GPT-5 scored 52.6 across seven tests, followed by GPT-5.5 at 55.9 and GPT-5.6 Sol at 56.1 below Claude and Gemini models.

On the current Text Arena leaderboard, the stronger writing-oriented variant GPT-5.6 Sol sits 9th overall in creative writign, while base GPT-5.5 sits 22th overall.

Technical specs

  • Context window (input): 1,050,000 tokens. Output window (single response): up to 128,000 tokens. OpenAI’s model page lists GPT-5.4 with a 1.05M context window and a 128K max output length, which is a major jump over the earlier GPT-5 generation and gives you enough room for long briefs, source packs, style guides, and outline scaffolds in one session.
  • Modalities: text and image input, text output. Audio and video are not supported on the model page. That is enough for most writing teams, especially if you are feeding screenshots, PDFs, charts, or visual references into the draft.
  • Availability / hosting: ChatGPT, OpenAI API, and Codex. OpenAI released GPT-5.4 across ChatGPT, the API, and Codex on March 5, 2026.
  • Private / local: No local or offline self-hosting. GPT-5.4 is delivered through OpenAI’s managed services rather than as downloadable weights.
  • Tools: GPT-5.4 supports web search, file search, image generation, code interpreter, hosted shell, computer use, MCP, and tool search in the Responses API. This is one of the reasons it is so useful for research-backed content and complex editorial workflows.

Pricing

API pricing: GPT-5.4 is listed at $2.50 per 1M input tokens, $0.25 per 1M cached input tokens, and $15 per 1M output tokens. For prompts above 272K input tokens, OpenAI says GPT-5.4 sessions are billed at 2x input and 1.5x output rates across standard, batch, and flex processing.

ChatGPT subscriptions: OpenAI’s pricing pages say GPT-5.4 access is included across paid ChatGPT plans, with GPT-5.4 Pro reserved for Pro and flexible access on Business and Enterprise. OpenAI’s plan announcements list ChatGPT Plus at $20/month and ChatGPT Pro at $200/month.

6. Grok (4.2)

Grok 4.2 remains is a distinctive LLMs for writing. It has more edge than most frontier models, reaches for less obvious references, and is less afraid of sounding opinionated.

That makes it especially interesting for essays, op-eds, founder-led content, and any piece where you want a sharper point of view instead of a neutral explainer. Claude is still more graceful and Gemini is more controlled, but Grok can produce the line or angle that the safer models avoid.

The benchmark puts that personality into context. Grok 4.2 scored 58.0 across all eight tasks. Its sentence rhythm and low GPT-2 proximity were relatively strong, but heavier prose, weaker syntax diversity, and limited section variation made the output look more model-like overall

In the current Chatbot Arena, Grok-4.20 is ranked 19 th in creative writing ranking.

Technical specs

  • Context window (input): 2,000,000 tokens. Both the reasoning and non-reasoning versions of Grok 4.20 Beta are listed by xAI with a 2M-token context window. That gives you enough space for full transcripts, long research notes, multiple source documents, and a style guide in one prompt.
  • Output window: xAI’s public model and pricing pages do not publish a separate max output-token cap for Grok 4.20 Beta. What they do publish is output pricing, and the API supports streaming responses. So the practical takeaway is that long responses are supported, but xAI does not currently expose a neat “max output window” number the way some vendors do.
  • Availability / hosting: xAI API and the Grok app / Grok.com ecosystem. xAI also says Grok 4 is available to SuperGrok and X Premium+ subscribers.
  • Private/local: There is no published local or offline self-hosted Grok release. xAI’s documented access paths are its own managed API, enterprise API, and consumer app subscriptions rather than open weights or downloadable local builds.
  • Speed: This is one of Grok 4.20 Beta’s strongest selling points. xAI explicitly markets it as “industry-leading speed,” and its broader API lineup includes Grok Fast and Grok 4.1 Fast variants when you want cheaper or higher-throughput iterations.
  • Data posture / vibe: Grok is tightly connected to xAI’s real-time search and tool stack. In practice, that means fresher references, more cultural texture, and sometimes a more aggressive tone than other models. That can be an advantage, but only if your editorial guardrails are in place.

Pricing

xAI API pricing:
Grok 4.20 Beta (reasoning): $2.00 / 1M input tokens and $6.00 / 1M output tokens.
Grok 4.20 Beta (non-reasoning): $2.00 / 1M input tokens and $6.00 / 1M output tokens. Both versions carry the same 2M-token context window.

Cheaper fast variants: xAI also lists Grok 4 Fast and Grok 4.1 Fast variants at $0.20 / 1M input and $0.50 / 1M output, also with 2M context. Those are the better choice when you want speed, summarization, or cheaper first drafts at scale.

Consumer access: xAI says Grok 4 is available to SuperGrok and X Premium+ users. On X’s public help page, Premium+ starts at $40/month or $395/year on web, with higher Grok limits included.

7. GLM (5.2, 5,1)

GLM 5.2 earns a place in this ranking because it performed better than several more familiar model families in the new benchmark. Its writing is not equally strong on every signal, but it can produce clean and efficient prose.

Its main strengths are sentence-length variation and relatively light prose. That gives the writing a less compressed feel than many frontier models. The tradeoff is that its token choices, syntax patterns, personal tone, and section-to-section movement can still look more predictable.

In the sylistic benchmark, GLM 5.1 scored 52.0 across six writing tests.

Its current Chatbot Arena placement for GLM 5.1 is 27th.

Technical specs

  • Model tested: GLM 5.1 (z-ai/glm-5.1).
  • Benchmark coverage: six of eight writing tasks.
  • Context, hosting, and local deployment: not supplied in the source data for this rewrite; verify against Z.ai’s current model documentation.

Pricing

The supplied benchmark does not include GLM pricing. Verify current API and subscription rates with Z.ai before adding a specific figure.

8. DeepSeek (V4 / V3 / R1)

DeepSeek remains one of the most balanced open-model families for writing. It performs well across essays, structured nonfiction, SEO, and reasoning-heavy drafts without being built around only one style of content.

The open-model route is a major advantage. You can host it privately, customize the workflow, and retain more control over data and model behavior. That makes DeepSeek especially appealing to writers and teams that value independence from closed ecosystems.

The stylistic benchmark is more cautious about its raw prose. DeepSeek V4 scored 62.1 across seven tasks. Sentence rhythm, prose density, and personal tone were its better signals, while token predictability, syntax diversity, and section variation made the output look more model-like. DeepSeek is still strong for structure and reasoning, but the final prose usually benefits from a dedicated voice edit.

In LMSYS Chatbot Arena, DeepSeek-V4 sits in the 38th rank across the creative writing category.

Technical specs

  • Context window (input):
    • DeepSeek-V3.2-Exp (API): 128K tokens. Max output: 4K by default, up to 8K on the non-thinking chat model; the thinking (“reasoner”) path supports 32K default / 64K max output.
    • DeepSeek-V3.1 (providers): commonly exposed with ~164K context; exact limits vary by host.
    • Open-weight checkpoints (V3 / R1 / V3.1): context depends on the build you run and the serving stack. Start from the model card and your inference framework’s limits.
  • Private/local LLM:
    • Yes (open-weight): You can self-host the V3 / R1 / V3.1 checkpoints for full data control and fine-tuning.
    • Managed options: R1 and V3 variants are also available on major clouds and model gateways if you want governance without running GPUs yourself.
  • Speed / responsiveness:
    • On the API, generation is competitive and benefits from prompt caching (see pricing below). On local runs, throughput depends on your quantization, batch size, and GPU VRAM. (DeepSeek publishes optimized kernels and infra repos that help you squeeze latency.) a
  • Modes you’ll use:
    • Chat (V3.x) for general writing with tool use.
    • Reasoner (R1 / V3.x “thinking” modes) for deeper planning and multi-step tasks. V3.1 adds a hybrid switch between thinking and non-thinking via chat template—handy when you only want “thinking” on hard passages.

Pricing

  • Official API pricing (V3.2-Exp):
    • Input: $0.28 per 1M tokens ($0.028 per 1M with cache hit)
    • Output: $0.42 per 1M tokens
    • Context: 128K; output caps noted above (8K chat; 64K reasoner).
      These rates are listed on DeepSeek’s own pricing page.
  • Third-party hosts: Some providers advertise ~164K context for V3.1 and similarly low prices (sometimes even lower promotional rates). Always validate context and limits per host before you budget.
  • Open-weight (self-hosted): Model files are free to download; your costs are compute + ops. If you already own GPUs (or rent spot instances), this can beat API pricing for heavy throughput—at the expense of MLOps overhead. Start with the official repos and kernels.

9. Qwen (3.8 Max / Qwen 3)

Qwen is a strong contender among open and openly deployable model families. I find it particularly useful for essays, literature reviews, multilingual work, and structured nonfiction.

The outputs are often detailed and well organized, though they can become over-formatted or too procedural. In creative writing, Qwen is capable, but it does not yet have the same narrative instinct as Claude or the same editorial sharpness as Grok.

Qwen 3.8 Max scored 58.0 across four benchmark tasks. Its token choices were impressively unpredictable, but sentence rhythm, syntax diversity, personal tone, and section variation were weaker. Because only half of the eight-task battery was completed, the result is directional rather than definitive.

On Chatbot arena creative writing category, Qwen 3.8 max achieves the 3th place.

Technical specs

  • Context window (input). There are two tracks to know:
    • Open-weight releases. Qwen shipped open models with up to 1M-token context (e.g., Qwen2.5-7B/14B-Instruct-1M), so you can self-host long-context drafting without a vendor lock.
    • Hosted/API SKUs. Specs vary by provider. On OpenRouter, Qwen-Max (based on Qwen 2.5) lists ~32K context. On Alibaba Cloud’s Model Studio, the Qwen Plus / 2.5 family offers tiered pricing up to 1M input tokens per request (with separate “thinking” vs “non-thinking” modes). Check which SKU you’re actually calling before you paste a megadoc.
  • Output window. Output caps follow the SKU (and “thinking”/reasoning mode). Alibaba documents separate prices/tiers for non-thinking vs thinking generations as context grows (≤128K, ≤256K, up to 1M).
  • Private / local. Qwen maintains official open-weight repos (Hugging Face, GitHub). You can self-host and fine-tune, or run managed via Alibaba Cloud. This is ideal if you want EU-style data posture or on-prem control.
  • Speed / modes. Expect faster, lighter SKUs for drafting and reasoning (“thinking”) modes for deeper planning. As with other MoE models, throughput depends on the host and your quantization/batch choices when self-hosting.

Pricing

Budget depends on where you run it.

  • Alibaba Cloud Model Studio (API). Qwen Plus/2.5 family uses tiered pricing by context band with separate rates for non-thinking and thinking generations. Example bands documented today: ≤128K, ≤256K, and (256K, 1M] inputs—with per-million token rates scaling at each tier. This is your reference if you’re deploying inside Alibaba Cloud.
  • OpenRouter (hosted). Qwen-Max lists around $1.60/M input and $6.40/M output with ~32K context (providers may vary). Good for quick pilots and prompt testing across vendors.
  • Self-hosted (open-weight). Model files are free; you pay compute + ops. If you already have GPUs (or rent spot instances), self-hosting the 1M-context open releases can be cost-effective at volume—provided you’re ready for MLOps work.

10. Muse 1.5 (Sudowrite)

Writing quality

Muse is the only LLM in this ranking purpose-built for fiction. Sudowrite’s narrative workflow helps it plan, revise, and stay close to the brief, so longer scenes and chapters tend to meander less than they can with general-purpose models.

Its strongest qualities are voice, dialogue, scene movement, and the willingness to write darker or more adult genres. The tradeoff is specialization: it is exclusive to Sudowrite, internal consistency can still fall behind Claude on intricate lore, and Gemini or ChatGPT remain better choices for factual, technical, or SEO-heavy work.

TRY OUT MUSE LLM FOR FREE HERE

Technical specs

  • Context window (input). Sudowrite doesn’t publish a token count for Muse. Practically, you work at the scene/chapter level inside Sudowrite’s Draft/Write tools, with streaming output and tools (Expand, Rewrite) that keep the active context focused. If you need hard token limits, Muse isn’t marketed that way; it’s positioned as a fiction workflow rather than a general API model.
  • Output window. Generations stream and can be extended (“continue”) inside Draft/Write. Muse 1.5 specifically advertises longer scenes and tighter instruction following compared with earlier builds.
  • Private/local. No self-hosting. Muse is exclusive to Sudowrite (web app). Sudowrite states your work isn’t used to train Muse and emphasizes an ethically consented fiction dataset. If your priority is author-friendly data promises inside a hosted tool, this is the draw.
  • Filters and tone. Muse is marketed as “most unfiltered” on Sudowrite (can handle adult themes/violence) and as actively de-cliché’d during training. Good to know if you write darker genres—and equally important if you need to keep drafts brand-safe.

Pricing

Sudowrite sells access to Muse via credit-based subscriptions (browser app). Tiers (monthly billing shown on the live pricing page):

  • Hobby & Student — $10/mo for ~225,000 credits/month.
  • Professional — $22/mo (page shows 450,000 and 1,000,000 credits copy; treat 1,000,000/month as the current headline included amount on the page).
  • Max — $44/mo for ~2,000,000 credits/month with 12-month rollover of unused credits.
    All tiers include the Sudowrite app; Muse is selectable as the default model in Draft/Write. Always confirm the current inclusions and yearly discounts on the pricing page before budgeting.
Free Prompts

Free Prompts and Ebook to Humanize Your Text

Download Now
Share
Buchert Jean-marc

Buchert Jean-marc

Confirmed AI content process expert. Through his methods, he has helped his clients generate LLM-based content that fit their editorial standards and audiences expectations.

All Posts

Watch more about it

Explore our latest videos

Related Articles

Explore our tips and prompting techniques for quality AI content

Claude Writing Full Review

Ask writers which LLM produces the best prose, and Claude often comes up. But what…

GPTHuman Review & Bypass Rate

GPTHuman is one of the newest tools in this space, yet it’s already getting a…

Sudowrite Review : Is It Really Worth it ?

Do you want to know if Sudowrite writing tool is good ? Here's our opinion