Show HN: Voyager – write a web crawler/scraper as a state machine in Rust
github.com
github.com
What are your impressions of scraper and html5ever? When I initially looked at HTML/XML parsing libraries for Rust, there didn't seem to be a standout library such as serde_json for JSON data. I was also considering using scraper + html5ever. However, I'm curious if scraper adds enough to warrant the additional dependency as opposed to directly using html5ever.
Once I add the tokio feature, they all run as expected.
Basically this a just a futures_timer::Delay [0] that is reset after each request, which is non blocking.
But you can also spawn all tasks upfront via `tokio::spawn`, and let them wait on a `tokio::sync::Semaphore` before making the request. The drawback of this is that you might allocate more memory for tasks upfront - but if you don't have an extremely high number it might not matter.
[1] https://github.com/quinn-rs/quinn/blob/de627437bc7d836564c36...
It was recently used to implement rate limiting middleware for both tide[1] and actix[2]
[0] https://docs.rs/governor/0.3.1/governor/_guide/index.html