Show HN: Automated writing analysis of Twitter and Reddit feeds
anthropologize.com
anthropologize.com
I had originally started the project to grab Flesch-Kincaid grade levels, but it was taking me 45 minutes to analyze about 20 pages of comments, so I switched it out for Automated Readability Index.
I really wanted to include HN, as I feel it's in a period of social shift, but it wouldn't have the same sort of statistical significance due to its size (it would most likely just be meaningless jarring jagged lines across the graphs). If someone can think of a good way to include HN, I'm all ears. I'd also be interested in other methods of writing analysis that I could include.
It might be even more interesting to do comparisons between subreddits and comparisons between hashtags or followers of certain figures.
Also do you have a link to the source?
Keep up the good stuff!
I have been thinking about subreddits and hashtags, but I'm not sure what would be interesting to people and still have enough statistical significance to update every half-hour/hour. If I were going to release this on Reddit, I'd have probably arranged it as a comparison between all the major subreddits. As for hashtags, their impermanence makes them a difficult measure.
I don't have the source up yet, but it's pretty simple. I just get the json output of /all/comments and get the Twitter streams for the words I'm checking with ntwitter and then do the basic math to analyze it every half hour. I render the page myself with the backlog of values and then send it updated info as I process it, so if you leave the page open, it should keep fresh.
One way to approach the "significance" problem for smaller communities like HN is to create larger bins of misspellings/correct spellings. However, I don't think you're going to see many "anyways" or "yea" on HN at all, much less significant fluctuations over time.
Finally, an interesting question might be to what extent individuals fluctuate in their ARI/FK levels in different contexts. What if a poster in /r/lolcats writes a very asinine comment but then goes five minutes later to /r/programming to speak intelligently. Or is it the case that an idiot is an idiot, regardless of context?
HN comments tend to be more varied and sparse, so I think you're right that measuring yea:yeah, etc. wouldn't be too enlightening. That said, I've noticed a defined qualitative shift in HN comments over the past few months, and I'd like to develop ways of measuring that before they reach Reddit/Twitter levels.
As for your last point, I could track individuals but a relational comparison based off of /all/ data would be pretty difficult due to the number of comments vs. the few number of any individual's comments. Also, ARI isn't a great metric (hence me putting it in a tiny graph) because it measures chars instead of syllables. For example, "FFFFFUUUUUUUUU" has the same score as "constructivism".
They don't, at least, seem in the same category as clear spelling errors like copyright/copywrite, or bureaucracy/beaurocracy.
I think you're measuring something other than spelling here, closer to measuring prevalence of certain dialects. Especially in the "anyways" example: if there's something objectionable about that word, it's a usage objection, not a spelling objection. People really do say "anyways", and that would be the correct way of writing it if you accept that usage. Spelling it "enyways" or "annyways" or something would be an orthographic error, on the other hand.
As for anyways, I think your classification there will depend on who you speak to (which is the point of the analysis).
As Jorge Luis Borges once said, all words were once neologisms.
For something to be standardized or make it into a dictionary, lexicographers look at its popularity. If we reduce language to the most simplest idea, it is exactly this: a set of popular phonetic sounds from which humans derive agreed similar meaning.
The phenomenon of adding an extra sound to the end of a word is called paragoge.
For anyone that is interested and is not familiar with it, linguistics has a lot to say on the addition and subtraction of sounds at the beginning, middle, and end of words. Here is a link: http://en.wikipedia.org/wiki/Metaplasm
Expanding on this could lead to some interesting insights. How about "could of" and "could have"? and the old "could care less"...
Edit: As for the axes, I made some stylistic decisions that I wouldn't dare to make were this being submitted to the Annual New Media Socio-Linguistics Journal.
Admittedly that doesn't apply here with "anyways" but it perhaps does with "yea". Perhaps throw out tweets over a certain length?