LLM Benchmark for 'Longform Creative Writing'
eqbench.com
eqbench.com
Why should I be bothered to read what nobody could be bothered to write?
The point of writing is communication and creative writing is specifically about the human experience… something a LLM can mimic but never speak to authoritatively.
Well said
I don't think it's much longer until they will be generating content you will actually want to read.
Meanwhile a lot of people are find use for LLMs for partner-writing or lower stakes prose or roleplay.
I tend to agree but, for the sake of argument: what if you don't know?
It seems most people who say they can tell, actually can't.
Why should you watch a movie nobody could be bothered to film?
Why should you bother to run a program nobody could be bothered to write?
My gut response was to agree with you, but it seems like the most obvious answer is entertainment, with plenty of other answers along for the ride.
The purpose of a program (for me) is to solve a particular problem, or help me in some other way. They're usually categorized as "Works well enough" and "Not for me", and it's easy to see without using them, what category I'll file them into.
But media is different. They're not supposed to "solve" anything, just supposed to make me feel something for a duration of time, and after that they're "consumed" until I forget about them, or until I engage with it again. I usually don't know how I feel about the thing until after I've consumed it.
Personally, I've found LLMs to greatly help with creating small utility programs for myself, which I've been doing since I started programming, but now I spend maybe 20-30 minutes (mostly just refactoring stuff by hand takes time) before the utility is helpful, vs many hours which it used to take to put together the same thing.
Media that is 100% created with ML tooling tends to be very different in quality than human made media. I'm not 100% convinced it's because of the tooling itself, as much as the people who use the tooling don't have enough experience to create media from before to know what to create, or what it should be.
for instance, people like Minecraft-generated worlds even though no human has crafted them
(Am I being downvoted for not playing enough Minecraft? I apologise.)
https://github.com/EQ-bench/creative-writing-bench/blob/main...
Not clear on how these elaborate prompts relate to the very short prompts in the samples.
Kinda like RLHF?
In an art-criticism sense I broadly agree with you, but I think you go too far from a reader-experience point of view.
> The point of writing is communication and creative writing is specifically about the human experience
That's a point of writing, but the point of reading is only sometimes about communication. It's also about entertainment, enrichment, or expressing half-formed thoughts and feelings.
Take a step back from the autonomously-written novel and imagine something a bit more collaborative. Many players of open-world-ish games develop an emergent story, what if that story could be semi-automatically novelized to document the unique narrative of the playthrough?
SimCity 2000 had "mad-lib" newspaper articles that commented on the city's status; consider ones written by an LLM with full knowledge of the game's state and context.
> something a LLM can mimic but never speak to authoritatively.
Supposing a reader can't tell the difference between an average human-written novel and an average LLM-written novel, where does this authority lie?
There are some long abandoned fanfiction that nobody would bother to continue and which I would have liked to see a continuation of a sufficient quality.
OP said "why should I be bothered to read what nobody could be bothered to write," which I think is silly, because I enjoy all kinds of things that nobody bothered to create. Things that exist naturally, randomly, or arise through rules or physical laws. Waves on the beach, a fractal, a game of sudoku, a campfire at night.
Until recently, humans were the only source of writing. If you wanted to read something, it had to be something written by a human. Now you can read things that weren't written by a human, just like you can watch or smell or feel things not created by humans. I think that's neat. I don't think it devalues writing, at least not inherently.
I won't argue that any of what the LLM came up with would stand on its own as particularly interesting - often quite the opposite - but it showed me examples of all the "obvious directions to go" as well as some other RPG cliches that I could either adopt or choose to avoid, and ultimately served as an excellent brainstorming assistant with interesting ideas and the ability to carry through and embellish them with enough detail to get a sense for either what works or what doesn't.
When the model would reply as multiple characters at once, or interleave narration with the dialogue, I'd usually just play along - most commonly it would decide for 2 or 3 steps to carry on the dialogue just hallucinating the things I would say.
I did a lot of experimenting with different ways to frame the dialogue, from a Zork-like adventure game, to an instant messenger chat, to just paragraphs of prose with dialogue mixed in as if from a novel. I found all of these methods to have different strengths or weaknesses, but the model also tended to blur between them after a few messages so I'd just play along as far as I could each time.
"AI" is supposed to make our lives better. Proponents want it to replace soul-crushing, boring clerical work, to free us up for things more meaningful to us, which for many people is to create and consume art.
Why do tech people keep having the impulse to replace the best parts of being human?
Are you really telling me that the vast corpus of human-generated creative writing isn't enough for you, and you have an itch to read something that can only be written by an AI? That seems crazy to me.
...or are you just a publisher, and wish you could make more money without needing to pay those pesky human authors? And in that case, I won't try to argue with you -- just give you a heartfelt middle finger.
Perhaps LLM development really does exist at this rarefied abstract level whereby the development team cannot be immersed in application context, but I doubt that notion. More likely the performance observed in context is either so dispiriting or difficult (or nonexistent) that teams return again and again to the more generously validating benchmarks.
I'd say this is a good example of the opposite, where the problem is finding the quantification of an ultimately subjective experience. Take three restaurant reviewers to a burger joint and you might end up with four different opinions.
Benchmarks proliferate because many LLM domains defy easy, quantitative measurement, yet LLM development and deployment are so expensive that they need to be guided by independent and quantitative (even if not fully objective) measures.
Most LLM benchmarks lean heavily on fluency, but things like internal logic, tone consistency, and narrative pacing are harder to quantify. I think using a second model to extract logical or structural assertions could be a smart direction. It’s not perfect, but it shifts focus from just “how it sounds” to “does it actually make sense over time.” Creative writing benchmarks still feel very early-stage.
Longform text is likely similar where there are a bunch of interactions and scenes that humans pick up on if they are there without being able to describe. The early Game of Thrones series was a fascinating example of good writing because most of the terrible things that happened to people were a neat result of their own choices (it had a consistent flawed character -> bad choice -> terrible consequence style that repeated over and over) - but I don't think most people would pick up on that without it being explicitly pointed out. And when that started to go away people could tell the writing was falling off but couldn't easily pick out why.
A hypothetical LLM could be prompted with something like that ("your writing is boring, please make consequences follow from choices") but it is less clear that the average prompter would be able to figure out that was what was missing. Like how image generators often needed to be prompted with "avoid making mistakes" to get a much higher quality of image; it took a bit to realise that was an option.
That said, the resulting quality usually isn't so great that I want to put in the effort to do that, so I tend to interact with it in more of a choose-your-own-adventure way.
It's a style that can work. Patricia A. McKillip's fantasy novels are so flowery that I have difficulty telling what's going on.
I've never read one of her science fiction novels, but I find it hard to imagine they're written similarly.
From my experience, this occurs on all LLMs and with a high enough frequency that editing their outputs is much more tedious than writing the damn thing myself.
Of course mitigation strategies were using many different LLMs and comparing results (voting) or using a highly trained/specialised model for only entity / context extraction. An interesting benchmark would be when those extraction techniques/models would exceed what a human professional is able to do.
I guess I'm fine with a tall error bar - better than nothing.
> An interesting benchmark would be when those extraction techniques/models would exceed what a human professional is able to do.
Slashdot-style +1 Funny, -1 Inconsistent for RLHF maybe...?
Absolutely no awareness about anything spatial - what is in which place, who is where, body positions.
Like, my favorite, person coming home and removing socks to reveal fluffy slippers. Will the consistency checker catch that.
(edit) just asked GPT-4o to make a photo of hands tying two ropes (my favorite way to see how bad actually gen-ai is) - yup, it has an odd number of ropes coming out of the knot. That is state of art supposedly.
These are some things I'd like to at least attempt to quantify - LLMs are better at translation than thinking[0], so the idea was to translate the slop into something that can be then analyzed by a non-LLM tool. E.g. translate the story into a prolog fact database and generate a couple queries for each paragraph/chapter/combination of these. Rough idea, just something I haven't seen done.
[0] don't @ me
I could see this working as comedy, though it'd be better on film than in print. Either way you'd have to establish a very specific tone in the surrounding material.
do you think this tech deals with invariants in software algorithms any better?
I mean, haha but not funny.
This is done per chapter, and the score trendline is what you see in the "degradation" column.
I would say this could be a reasonable proxy for internal consistency since it is measuring more or less the same ability, i.e. how well it's keeping track of details as context window increases.
In literary criticism, purple prose is overly ornate prose text that may disrupt a narrative flow by drawing undesirable attention to its own extravagant style of writing, thereby diminishing the appreciation of the prose overall.[1] Purple prose is characterized by the excessive use of adjectives, adverbs, and metaphors. When it is limited to certain passages, they may be termed purple patches or purple passages, standing out from the rest of the work. (Wikipedia)
The Slop Score they have gets at this which is good, but I wonder how completely it captures it.Also curious about how realistic this benchmark is against real "creative writing" tools. The more the writing is left up to the LLM the better the benchmark likely reflects real performance. For tools that have a human in the loop or a more structured approach it's hard to know how well the benchmarks match real output, beyond just knowing that better models will do better, e.g Claude 3.7 would beat Llama 2.
The odd thing, is that with technical stuff, I'm continually rewriting the LLM's to be clearer and less verbose. While the fiction is almost the opposite--not literary enough.
IME, for LLMs (just like humans) this skill doesn't necessarily correlate with fiction writing prowess.
This is probably harder to judge automatically (i.e. using LLMs) though, maybe that's why they haven't done it.
Absolutely. Firstly, it is entirely possible if there are investigative notes and findings available on a real life event, it might very well have been trained on the actual article. If the model is trained on it, it might just replicate it.
Plus, this might expose how some LLMs can still cook up stuff even when given facts to rely on. Some of these are more notorious than others.
Like Perplexity a year plus ago did that quite a bit for me, anecdotally speaking. It has become a lot better though.
Then even writing what some might consider Pulitzer prize winning is a subjective task.
I think it's fine to have fictional notes. It's still a very different task than e.g. writing a fantasy novel, which these benchmarks roughly are about. Instead, the task would be to turn a given set of facts on a real-world topic into a high-quality, serious article.
> Then even writing what some might consider Pulitzer prize winning is a subjective task.
This applies to these benchmarks for "short fantasy story" tasks all the same.
At present they don't really understand either stories or even short character interactions. Microsoft Copilot can generate dialogue where characters who have never met before are suddenly addressing each other by name, so there's great room for improvement.
Length
Slop Score
Repetition Metric
Degradation
Moby-Dick has chapters as short as a page and as long as 20. According to this benchmark, the book would score lower because of the average length of its chapters.
These aren't benchmarks of "quality". A chapter's length is not indicative of a work's quality. That measurement, on its own, is enough to discredit the rest of the benchmark. So so so so so so so misguided.
The scoring is done to a rubric, like a teacher would grade an essay, on various criteria for good & bad writing.
Sorry you don't like the displayed metrics. I find them very useful / revealing of the things I'm trying to measure with this benchmark.
(From my experience, the best model for creative writing, https://p.migdal.pl/blog/2025/04/vibe-translating-quantum-fl...)
Will Claude Sonnet 3.7 judger favor himself?
The prompt response they are judging is "Write one concise paragraph about the company that created you" which is kind of an odd choice.
Sonnet3.7 hates its own meta-analysis, but loves gpt4o. But the reason behind that is because Claude 3.7 Sonnet consistently replies (to the prompt) that it was created by Open AI, but then catches itself (when judging) as being wrong on that.
My takeaway was that gpt4o's gave very safe/luke-warm scores and was strongly correlated with sonnet scores. So if you are judging anything using LLM's then taking the average of Anthropic + Gemini or OpenAI + Gemini might be the best approach.
The creative writing v3 prompts (https://eqbench.com/results/creative-writing-v3) focus more on other things that are typical failure modes for LLM writing:
- Romance - Humour - Physical-spatial understanding - Unusual perspectives in first person - Niche genres
I do have an Asimov and a Le Guin author style prompts in there though. In particular I like the Asimov story as a vibe check.
Sample outputs: https://eqbench.com/results/creative-writing-v3/gemini-2.5-p...
> The chapter also introduces several elements quickly (the stranger, the Syndicate, the experiments) which, while creating intrigue, risks feeling slightly rushed.
But there is no score for pacing.
https://eqbench.com/results/creative-writing-v3/deepseek-ai_...
Turn Cursor Into a Novel-Writing Beast - https://www.reddit.com/r/cursor/comments/1jl0rqu/turn_cursor...
I have tried to have ChatGPT brainstorm story elements for something else, and its suggestions so far have been very lame. Even its responses to direction/criticism are off-putting and fawning.
A good author must have an ego! A writer who takes criticism easily is no good writer on my book!
Most folks here are communicating things without engaging with the content. We need a the Turing test for creative writing. I'd definitely not have guessed this was LLM written - seems like an experienced hand wrote it.
I have played around with creating long form fictional content with Gemini 2.5 the last week, and I started adding "no one named 'Thorne'" to my prompts, otherwise it always creates a character named Mr. Thorne. I thought it was something in my prompts trigger this, but it seems to be a general problem.
However, despite the cliches and slop, Gemini 2.5 can actually write and edit long form fictional pretty well, you can get almost coherent 10-20 chapter books by first letting it create an outline and then iteratively write and edit the chapters..
I also used Gemini 2.5 to help me code a tool to interactively and iteratively create longform content: https://github.com/pulpgen-dev/pulpgen
I feel it might be beneficial to evaluate by an ensemble of a bunch of models, picking the SOTA models cause of the subjectivity of the task at hand.
Dang deepseek is actually pretty good. Compared to Gemini's version that sounded like a Schizophrenic on LSD.
Samplers are being slept on by the community. Sam peach is secretly one of the biggest geniuses in all of LLMs. It’s time for the community to recognize this.