Show HN: Getting started with Puppeteer and Chrome Headless for Web Scraping
github.com
github.com
For example with Puppeteer you can do page.injectFile("jquery-3.2.1.min.js"). I think that would simplify your evaluate() calls.
It would also be easy to speed up the whole process by doing a single evaluate() call per page with all your scraping code in it.
BTW we just released an article with tips & tricks for Headless Chrome: https://blog.phantombuster.com/web-scraping-in-2017-headless... What do you think?
Correct me if I'm wrong, but if I'm notm mistaken Selenium IDE has been discontinued due to lack of mantainers, and that has little if any relation to Chrome Headless.
The IDE is just a more effective way of programming test behavior; the Selenium webdriver is still up and working with straight code (as is the case of this tutorial).
We switched to chrome headless after a post from thoughtbot made me question Capybara-WebKit's future.
see https://seleniumhq.wordpress.com/2017/08/09/firefox-55-and-s... and associated HN discussion https://news.ycombinator.com/item?id=15061605
I have updated the article so it doesn't seem like the two events are related.
""Somewhat similar is the case with Internet that we traversed today in quest of data.""
1. Please do not test a web app with Chrome only, we don't want to go back to a world with a single browser
2. > So, until puppeteer supports this, we will rely on jsdom, a package available via npm
JSDOM is not just a package on npm, it's an engineering piece of art
I hope Puppeteer become a standard.
I guess I can cheat by intercepting the request and respond with the html I already have. But I wonder if there is already something existing.
Initial assumption when reading the thread was that navigating to a data URI would be handled like entry of a data URI into the omnibox and still be allowed.
A small test case confirms that assumption - it works.
I got able to make it run inside a docker.
In this exact moment the example at the repo is just returning a blank PDF but the problem is at the API Gateway.
I have a repo outlining the basics here: https://github.com/jawj/web-scraping-for-researchers