The TechCrunch Bubble Index: Parsing Headlines to Quantify Startup Hype
toddwschneider.com
toddwschneider.com
article.published_at = Time.zone.parse(elm.css(".byline time").first.to_h["datetime"])
# timezone seems inconsistent, no big deal because we only care about date anyway
I've written a TechCrunch scraper myself, and it turns out the reason the time-zone is inconsistent is TechCrunch outputs the time in 12-hour AM/PM...but forgets to include the AM/PM. So the hours just loop around from 0 to 12.I haven't been back, but do catch an article or two when it's linked from HN. Does anyone here still frequent TC enough to comment on the current editorial direction?
I think it would have been better to keep this pretty powerful heuristic for yourself :)
http://blog.itrendcorporation.com/2012/07/23/no-coverage-for...
Your data is very interesting. Any change you could also "group by" company and list companies which are more frequently (repeatedly) covered by TC? I have a theory about that, would be interesting to test.
This is an amazing analysis. Thank you. I really enjoyed the aggregated "x for y" headlines.
I imagine that most of the funding articles are short. It would be interesting to see how much of the total words were about fundraising (i.e. not just headlines, but the entire article).
http://i.imgur.com/Xqhahjs.png
EDIT: reduced unintentional financial verbiage.
For anyone equally confused, minimaxir means "Share this on Facebook" shares.