Dead simple Python crawler for extracting structured data from a website to CSV
blog.webhose.io
blog.webhose.io
I convert HTML to CSV regularly (multiple times a day). But I do not use Python; I use C. Actually I use flex to make filters which I compile as static binaries that read from stdin. This is in fact how I read HN. The HTML is converted to CSV and then the CSV is imported into a database.
Prior to using flex I primarily used sed. For many sites I still do; it's faster than having to compile, test, recompile.
If anyone has a website they want in CSV, and need something faster than Python or Ruby, just post the url. I like to think I am reasonably good at this, but I only do it for personal use on sites I'm interested in so who knows. For me HTML conversion to CSV and plain text is an art - I practice it every day.
Out of curiosity, why?
2. Turn unstructured, difficult to parse data adorned with HTML, and other window dressing into structured data that is easier to parse.
Python and Ruby have their places, I wouldnt say one is clearly better than the other.
The except: and continues is enough to cause a heartattack
Ruby's community is very pragmatic and Python's is more scientific. Most people I'm discussing learning programming with are interested in making something, and pragmatism is far better suited to their goals.
e.g. from my perspective, as someone who uses programming as a tool, writing my own scraper is boring (other scrapers work well enough for me). From my friend's perspective, writing a scraper is a fun experiment.