Sorry to hear about the issues you ran into when trying out Diffbot! If you have some examples, I'd like to take a look into whether we can improve.
In cases where our automatic extraction using computer vision isn't 100% accurate, we offer a visual interface for actually overriding our default extraction. This input is then used as additional training data for the ML models.
Mozenda et al. are great if you only need data from a couple of sites and you don't mind spending time manually specifying and maintaining CSS selectors for each website and page layout you need data from.
Our crawling, and proxy support, is fairly robust thanks to our hiring the creator of Gigablast [https://gigaom.com/2013/09/10/diffbot-brings-big-time-search...].
If you'd like to give Diffbot another go or you have some examples where the extraction could be improved, please let me know!