https://www.quora.com/What-is-the-best-solution-for-an-autom...
However, I tried it out and found diffbot failed for a lot of websites I wanted to crawl. Seems to work well enough for pulling meta data like title of the page, blog post titles etc, but anything beyond that it struggled. And since I don't know what the algorithm is doing or have any insight to how it works.
I rather much have a tool where I can have the accuracy and control and without having to write a full web scraper and running on an actual browser, not phantomjs or Webkit derivation where I have no visual feed back of what the scraper is doing. Like often, I want a live video of what the crawler is doing. What is it clicking on? Where is the website slowing down? Log file feels limiting.
Also, It's more than just crawling the website now. There's a lot of single page apps and edge cases where I find majority of tools fail. Requiring clicking on angular.js links, rendering javascript or websites requiring extension to be installed, needing to login to a website that redirects domain multiple times, trying out every single permutation of dropdown, checkbox, keywords in input form and crawling the search results with infinite scroll...some websites even give you bogus data because they can detect you are not using a real browser.
90~95% accuracy means jack all if 5% represents 99% of the value to you. It's actually not that hard to achieve an automated solution but it's when it fails for that one website you rely on for most of your data extraction needs, then it becomes a question of how fast can you debug and fix the situation....tough to do when you don't have the means to do it yourself.
https://www.cs.uic.edu/~liub/publications/kdd2003-dataRecord...
on the computer vision related front:
http://ir.lib.uwo.ca/cgi/viewcontent.cgi?article=3296&contex...
http://repository.cmu.edu/cgi/viewcontent.cgi?article=1045&c...
here's a good 'snapshot' overview of where we are with Automated data extraction:
http://repository.cmu.edu/cgi/viewcontent.cgi?article=1045&c...
and surprise surprise guess who has the best research on this front (public), it's Microsoft!
http://research.microsoft.com/en-us/um/people/sumitg/pubs/pl...
here's there computer vision approach called ViDE a while ago but was abandoned (?) no news since then:
http://research.microsoft.com/en-us/people/jrwen/vips_techni...
Most scraping/crawling tools are super easy to use like Mozenda and you can outsource it on freelancer to get people to create the crawlers for you. However, I soon gave up because mozenda charges like $0.10 CAD cents for each page downloaded (intermediate pages too) due to using their anonymous proxies and I ended up just giving up. Would be great if somebody like Mozenda offered an unmetered plan and preferably with some type of rotating proxies so I won't get throttled. I can imagine that diffbot is probably not using anything like that and running it straight from AWS which anyone who wants to block scrapers are able to add their public ip range to ufw. I'm finding it more and more difficult to scrape from AWS and other free scraping tools as website owners seem to block the AWS ip range mostly.