That's awesome, great work from the ScrapingHub team. We had a similar approach at SerpApi. We write open source rspec tests and run them daily.
e.g.:
- https://github.com/serpapi/test-knowledge-graph-desktop/tree...
- https://travis-ci.org/serpapi/test-knowledge-graph-desktop
- https://github.com/serpapi/test-organic-results-desktop/tree...
- https://travis-ci.org/serpapi/test-organic-results-desktop
- https://github.com/serpapi/test-news-results-desktop/blob/ma...
- https://travis-ci.org/serpapi/test-news-results-desktop
Producing reliable scrapers and parsers is very hard. Testing as much possible is the only way to go. Smart use also of JSON schema on Spidermon.