Hacker News' Reading Level
google.com
google.com
http://en.wikipedia.org/wiki/Flesch%E2%80%93Kincaid_readabil...
A brief search on wikipedia reveals a few readability tests, but they all seem to be based on sentence/syllable ratios, not content complexity.
http://en.wikipedia.org/wiki/Category:Readability_tests
And in general they all rank multi-syllable (longer) words higher. Which would mean a conversation between two Java API writers would be ranked higher than a ruby conversation :)
Java vs Ruby vs Lisp: http://i.imgur.com/tq3pA.png
Yes, It seems to be since it is only available for content in english.
Although, as far as I know they have not said anything regarding the algorithm used yet.
You can see (what it seems to be) the official announcement here: http://www.google.com/support/forum/p/Web%20Search/thread?ti...
What you are proposing is a statistically generated version of the Gunning Fog Index (http://en.wikipedia.org/wiki/Gunning_fog_index) or the Flesch–Kincaid test (http://en.wikipedia.org/wiki/Flesch%E2%80%93Kincaid_readabil...).
If I were Google, I'd try that, but I'd also try something like working out percentage deviation from a Markov chain generated from their crawl. A method like that would show that my first sentence is pretty unreadable, while an algorithm based on word complexity would see it as pretty simple.
Indeed, that was my second thought, but I wonder if the gains are really all that large over a raw statistical analysis of the word bag, and whether they're worth the extra analysis space/time. It really depends on what Google is planning on doing with this metadata, internally; if an order-of-ten precision is fine (to pick out decisive categorizations), the raw analysis may be all that's needed.
Holy Shnikes, that's a tough one to parse! I wasn't sure what you were saying here, so I'm gonna break it down, working from the end of the sentence:
1. I'd be lying if I said I don't doubt you are not incorrect
2. I'd be lying if I said I don't doubt you are CORRECT
3. I'd be lying if I said I don't think you are INCORRECT
4. I'd be lying if I said I think you are CORRECT
5. I think you are INCORRECT
The idea is that each of the previous statements are saying basically the same thing; I'm just cancelling negatives each time. Anyway, am I correct to assume that you think that the GP is incorrect?
I'd be lying if I said I don't doubt you are correct
I'd be lying if I said I think you are correct
Yes, I think the GP is incorrect.
https://encrypted.google.com/search?hl=en&tbs=rl%3A1&...
I was expecting a higher proportion in the "advanced" category.
For comparison:
Math Overflow: https://encrypted.google.com/search?hl=en&tbs=rl%3A1&...
OnStartups: https://encrypted.google.com/search?hl=en&tbs=rl%3A1&...
English Language & Usage: https://encrypted.google.com/search?hl=en&tbs=rl%3A1&...
CS Theory: https://encrypted.google.com/search?hl=en&tbs=rl%3A1&...
Seasoned Advice (Cooking): https://encrypted.google.com/search?hl=en&tbs=rl%3A1&...
Physics: https://encrypted.google.com/search?hl=en&tbs=rl%3A1&...
It's no surprise really, but it seems like the more technical jargon used on a site, the higher the reading level Google assigns.
So much for Simple English wikipedia!
http://www.google.com/search?q=site:news.ycombinator.com&...
A few examples that I thought were neat:
- msnbc (44/55/1) vs bbc.co.uk (15/82/2)
- facebook (40/37/22) vs linkedin (2/90/6) (I wondering why facebook has so much "advanced" content according to Google)
- wordpress (35/47/16) vs xanga (76/23/1)
- boston college (5/41/53) vs harvard (2/6/91)
I've published my findings, a complete ranking, and source code: http://log.largevoid.com/2010/12/ranking-colleges-by-reading...
Put simply, high reading level doesn't mean good content. (trying to prove your point :D)
Is it possible this phrase is an example of itself? :-)
I am sure it was indented as such.If I'm reading something I would prefer it be written as simply as possible. From a writers standpoint though this can be much harder than just explaining an idea in complex terms.
To paraphrase Pascal, "I apologize for the length of this letter, but I did not have time to make it shorter." I like to think length in this case meant complexity.
I mean, even if you don't care about losing people who don't have time to learn a more precise-but-obscure synonym, there's always people who learned English as a third language. Why alienate them just to look a little clever?
Even this comment has too many big words. If I had more time, I'd edit it down more. "alienate" would become "turn off" and "precise-but-obscure synonym" would become... I dunno, something simpler. Big obscure words that lock people out is really missing the forest from the trees.