748 karma · joined October 31, 2011
Software engineer | Founder & CEO @ Apify
https://x.com/jancurn
https://apify.com/jancurn
https://www.linkedin.com/in/jancurn
- at the moment we don't store the HTML content of the visited pages (except of the last one), so the only way to determine if something changed is to run the 'pageFunction' on each page again and compare the results. This can be optimized in certain situation, e.g. you can crawl a product listing and only go to product details page if some basic property changed. Saving HTML for each page is certainly possible, but after the crawler finished loading a page, running a pageFunction adds very little extra overheads.
- if a page cannot be loaded for any reason, a detailed description of the error will be present in the JSON results. We want to implement a limited number of retries for these pages, for situations the error is just temporary.
- certainly, if your crawling strategy cannot be expressed using simple pseudo-URLs, you can use the low-level 'interceptRequest' function to control exactly how each new page navigation request is handled (enqueued/ignored), tell the crawler which URLs refer to same pages and shouldn't be visited again etc. You can also enqueue arbitrary pages to crawl using 'context.enqueuePage'. In fact, you don't need to use pseudo-URLs at all and control everything from your code.
But look at Google Groups for example - there is an infinite scroll to get all the topics, posts are also loaded dynamically, so you have to wait some time to get them.
In the SFO flights example you have to deal with pagination also using JavaScript.
We wanted to build a powerful tool which can crawl and scrape almost any website out there. It's slower, but you can use bench of our nodes to do it in parallel.
- implement a mechanism to find and click active page elements, track the browser actions
- recompile PhantomJS to support POST requests
- implement a page queue with checks for duplicate URLs
- implement some parallelization and failover mechanism for case PhantomJS crashes (it does)
- possibly implement support for infinite scroll
- setup your own pool of proxy servers
- setup a database to store the results
and finally make the whole thing simple to setup and use :)
Please have a look at the service, play with the examples and maybe set up your own crawl. My co-founder jakubbalada and myself will be around here to answer your questions. We'd love to hear what you guys think!
I have no idea what did they do with the phone, did they read my emails? check my photos? installed some spyware? I have no idea, but I'm sure if I rejected to unlock it, I would be missed my flight. I felt like crap after that.
Interestingly, they didn't care about my second phone, an old Nokia, probably iPhone is more popular.