Visualizing Tolkien
5013.es
5013.es
I think the author took this way to literally. Granted, it inspired some fun (albeit odd and basic) data analysis. But the point of the reviewer isn't about word counts but about style and pace. That, narratively, the Silmarillion just felt like it continued on and on without clear sections of rising and falling action or significant breaks. Sentences were long and atmospheric, rather than short, quick and active.
Interestingly, I think the larger frequency of 'of' actually suggests this. I suspect the Silmarillion (I haven't read it in ... a decade maybe?) has a far more complicated web of relationships–not just between characters, but between places; at the very least, it concerns itself with genealogies more.
Why didn't the author try the most obvious test of reading difficulty of all: Flesch-Kincaid? [1]
[1] http://en.wikipedia.org/wiki/Flesch%E2%80%93Kincaid_readabil...
First, it's she, not he.
Second, because I hadn't heard about that obvious test at all back then. I never pretended to do a superserious scientific analysis but rather answer to the questions that came to my mind, by using a computer to validate hypothesis.
But many thanks for the pointers & suggestions, though! I was thinking about rebuilding this to make it realtime+interactive so that more than one text could be analysed/visualised, so I'll make sure to introduce the new tests.
I think I may have misunderstood the purpose of your article: I first read it as a serious attempt to use textual analysis to do a comparison of the comprehensibility of Tolkein's best-known works, in which case you fell way short of the mark by going no further than word counting. I think it started off like that, but on re-reading it I see that you say you hit a brick wall (paraphrasing) and decided to have some fun visualising your results so far; something you did very nicely. So perhaps I got the wrong end of the stick.
I thought about what you've taken on here. Doing things at the word level is pretty easy. Taking into account grammatical structure to get things like sentence lengths and clauses per sentence, or breaking words up into syllables (needed for F-K, for example) is considerably harder. I'm interested to hear what you come up with. Flesch & Flesch-Kincaid are US inventions so perhaps not obvious to everyone.
On he/she: I agonised over that for a good ten minutes over my breakfast. I read around your blog and your twitter page this morning for clues as to the appropriate pronoun. In the end I couldn't tell, so I went with 'he' because I stereotyped you. I nearly changed my wording to use constructions like "the author," "they," and various other mealy-mouthed alternatives but they were too ugly. So I did try, but I got it wrong. I apologise.
Flesch Reading Ease score, which is what I actually meant, does the same but with different coefficients to come up with a more granular difficulty score, usually in the range of 30-100.
They're both pretty arbitrary. The more I read up on this subject the more respect I have for the author's own attempts at an originality score. It's all subjective ultimately.
While I consider defaulting to 'he' to be legitimate and acceptable, I actually prefer the zie/zir gender-neutral pronouns when I think about it. http://santiago.mapache.org/nonfiction/essays/zie.html If I ever begin to agonize about the gender of the person I'm talking about, that's enough to kick me over into using GNPs.
The idea makes sense, but I'm unclear on how to actually measure it in a way that's normalized by page count.
Some suggestions: try removing stopwords (the, and, etc.), it'll bring out more variation on the analysis. Particularly in the circular graphs.
Try unique n-gram analysis, I suspect that will show something interesting. At the very least 2-gram (bigram/digram) analysis might show something cool.
Some other interesting comparative measures, since the author already has all of the tokens, try the Jaccard index between book pairs to look for similarity. (may even want to break the books down into major sections: e.g. LOTR can be viewed as both 3 and 6 books.
Some others to try, sentence length, word length, distribution of tf-idf scores, etc. etc.
fun fun!
And thanks for all the suggestions! I've made a note of them all. Good to hear from people who know more about the topic than me.
Just think about it - battles with hordes of orks, hero elves AND 30 or so Balrogs. And loosing / winning a battle doesn't just blow up a tower, it creates spasms in damn middle earth. The ring wars got nothing on that..
Personally, I loved the book, but I was hungry for all the "lore" about Middle-earth and its history that I could find. I think most people would really prefer a self-contained story.
Oh, and for the record, an underlying reason for many people disliking The Silmarillion may be that Tolkien never really finished it. I have a little bit of a writeup of the story behind that on my Tolkien FAQ site: http://tolkien.slimy.com/faq/External.html#SilmChanges
It's great. And I don't say that about any audiobook (and I listen to audiobooks quite often). The narrator does a great, great job.
Mapping semantic associations would be more interesting. Something akin to http://xkcd.com/657/ generated by the in-text juxtaposition of names & related verbs.
Out of curiosity, where did the author get a copy of the text to analyze?
"Turin" is the only name in The Silmarillion graph? Interesting.
I got them in txt files, online. I own the original books too. I would have typed them in if I had all the time of the world, of course.
(I'm the author)
So, in order to test the hypothesis "Silmarillion is harder to read because it has lots of stop words", you need to calculate the relative frequencies of lots of other texts and see if there's something special about Silmarillion's top 10 versus all other's.
Surely, you have already done that using LOTR and The_Hobbit, but a much bigger sample is needed. At the very least, you may want to use 10-15 other works of fantasy from different authors, and that will be just like a back-of-the-envelop test to see if it is worth to pursue this experiment with a statistically significant sample.
[edit] 1. Provided it is sufficiently large.
;)
Incidentally, I never read anything by Lovecraft, but am a big fan of Tolkien. I came up with this nickname on the spot.
Still, I have mixed feelings regarding your comment. Firstly, the author shows a true hacker spirit by trying to use the computer to solve problems, on the other hand, the author shows great courage to publish his/her thoughts on the Internet.
I do not want to live in a world where people are afraid of publishing their research (as insignificant as it might be) because of other people calling it "a waste of time" or "useless and pointless".