Colly – Scraping Framework for Golang
github.com
github.com
This is an honest question, I'm not trying to take a dig at anyone in particular.
Why? Because Amazon's search is outright broken. The number of results changes when you change sorting mode, and sometimes, sorting by a different criteria will just serve you a "no products found" error page.
I'll generally write a personal product comparison program when it becomes clear that I can't be certain that I can find the best product by hand. Often, even specialized websites that should have parametrized product search/filtering don't have their data properly indexed, so you have to scrape and parse it yourself. Another reason is to cross-reference with data from other sources. E.g. what laptop can I buy that has the best single-threaded CPU performance (within some other restrictions) [1]?
[0]: https://github.com/CyberShadow/choose-product/blob/master/am...
[1]: https://github.com/CyberShadow/choose-product/blob/master/le...
The solutions I write using homegrown utilities are both more elegant and faster than any framework or library I have ever seen posted to HN or recommended elsewhere. Not to mention smaller and more agile. IMHO.
All frameworks and libraries that I have seen will all fail given the right input i.e. fuzzing.
I think parsing and transforming the content from webpages is just viewed as work that no one wants to do because, for whatever reason, webpages are still unpredictable.
the bottle neck in scraping is never the parsing/DOM representation/traversal.
Bandwidth and IP limits are the most common bottle necks, but these can be solved using multiple proxies and ssh tunnels. Colly has built in support for switching proxies [1].
>but these can be solved using multiple proxies and ssh tunnels. Colly has built in support for switching proxies
interesting
>your server has limited resources.
possibly.
Edit - not picking on you, but given the quality and ecosystem of libraries and ancillary tools for scrapy, I don't even consider alternatives at this point. Good on anyone who does it to learn but for actual workloads I won't consider anything else.
Make sure to use DNS caching on the box else add it in Go.
Colly only supports a single machine via map of visited URL's. Would be great if you replace with a queue like redis or beanastalkd.
visitedURLs map[uint64]boolA sane approach to this is for example to create a separate file for each type (Collector, HTMlElement, Request, Response, ...) and its attached functions/methods.
SQLite has a different opinion — https://www.sqlite.org/amalgamation.html
a) it's automatically generated, so you can do dev work on the "easier" split version
b) Go comes with package management that helps solve the deployment issue
c) Question: Does the concept of a "compilation unit" affect the Go compiler the same way?
Um, how does this disagree?