The emotional arcs of stories are dominated by six basic shapes
arxiv.org
arxiv.org
Length 1: rise, fall
Length 2: rise-fall, fall-rise
Length 3: rise-fall-rise, fall-rise-fall (must be interchanging because for example fall-fall-rise would probably just be considered fall-rise)
In the Harry Potter example you can see that a higher frequency is very significant. But the exact frequency is probably rather arbitrary from book to book, whether there are ten peaks or five, say. So I'd imagine, over lots of texts, any particular higher frequencies is less significant, leaving the lowest modes to dominate, as you point out.
So overall, a rather uninspiring result, I felt.
Though if we're both missing something, it would be good to know!
Let me know if you are interested in investigating this properly.
non video description https://www.washingtonpost.com/news/wonk/wp/2015/02/09/kurt-...
the talk is funny though, i saw him give a much longer version.
On the other hand, he tells of showing a graph to his son, of a book his son had read, and being told the graph was completely wrong for part of the book---it turns out that part of the book was written from the antagonist's viewpoint and the "emotional valence" was apparently inverted.
See one of the follow-ups to http://www.matthewjockers.net/2015/02/02/syuzhet/.
The antagonist's viewpoint issue appeared in footnote 1 of http://www.matthewjockers.net/2015/02/25/the-rest-of-the-sto... and it does seem like the sentiment analysis would clearly be backwards overall if this kind of thing continued for most of a book. (An example might be if an author depicted people enjoying themselves while committing horrible acts, and used more individual words related to the enjoyment than to the acts.)
Thanks for the references.
Now I'm kind of curious about sentiment analysis of something like Fight Club.
http://channel101.wikia.com/wiki/Story_Structure_101:_Super_...
I wanted to also share https://en.wikipedia.org/wiki/Monomyth Joseph Campbell's very similar work that I think Dan gives credit to somewhere in that 101 series.
I couldn't tell if this could even identify the stories shape from their sentiment graph -- e.g. if you fed it "Cinderella" could it identify the "Rags to Riches" plot that they identify?
It would be neat to see if you could extract and identify plot points -- deaths, fights, breakups, betrayals, (maybe with some of those newfangled neural network thinggies) and look for patterns there.
Still, it's encouraging to see people working on this. Maybe the results seem basic because this field is still unstudied. I expect software tools for storytelling to change a lot in the coming years with all the exciting new work in natural language processing and machine learning.
I haven't read this paper yet (this post is falling off HN too fast!) but his technique was to use sentence-level sentiment analysis to get a time-varying signal, rub a Fourier transform against it, cut off all but the lowest frequencies, and returned it to the time domain to draw pretty pictures. He, too, came up with six basic arcs, I think, probably for the same reason that klue07 mentions but using statistics against a bunch of books.
There was a certain amount of press coverage at the time. The R package, syuzhet, is available on github[2]. Also, you can look at my notes on playing wiht syuzhet and R[3].
[1] Starting with http://www.matthewjockers.net/2015/02/02/syuzhet/
[2] https://github.com/mjockers/syuzhet
[3] http://maniagnosis.crsr.net/2015/08/exploring-syuzhet.html http://maniagnosis.crsr.net/2015/08/syuzhet-prodding-frequen...
See the part about "Generate Dramatic Game Pacing" http://www.valvesoftware.com/publications/2009/ai_systems_of...
Introduce likeable character - get them stuck in trouble - get them unstuck.