Web crawling and downloading ebooks with phantomJS
debuggerstepthrough.blogspot.co.il
debuggerstepthrough.blogspot.co.il
But that doesn't seem to be the case here. With Python I would have used a parser like lxml or BeautifulSoup (and I'm sure there is something comparable for JS) coupled with Requests async methods. That would probably not only end up with shorter and more concise code, but also be a lot faster.
One advantage is that it's not always instantly obvious if you'll need JS to execute before you can scrape a page. If you start out with a simply html parser and then find out that you needed the JS to run first, you're going to have to start over. If you start out using phantomJS and then find out that you don't need any of the original JS to run, your script still works.
I guess phantomjs is a good a tool as any, but there is really no need to evaluate Javascript for a bit of plain HTTP+HTML parsing.
> return [ $($('h2 a')[0]).attr('title'), $($('h2 a')[1]).attr('title') ];
which are 2 css selectors and picking an element. That is pretty much covered by all of the available http parser libraries
I guess that means that each e-book only has to be sold once in your country? The U.S. laws might be overly protective of IP, but that's an interesting problem for publishers who in theory need to earn a profit if they're going to continue as entities as well as for authors who need to feed their families.
This obviously wasn't a problem when the books were printed on dead trees, because you'd only share copies that had been purchased, and if you're friend was reading your book you no longer had access to it. Curiously, I could rent my copy of a book to you in the U.S. without violating copyright laws.
No.
It means that you can read a book and, if you really like it, you can buy it. It means that you can discover new authors, topics and so without a huge investment.
This can sound demagogic: I have never had enough money to buy the books I wanted, nor to waste it trying to discover new books and topics. But with downloaded books I learned about tech and other fields. Eventually, I bough more books (a lot from U.S) than if I had not discovered these topics.
Allowing private sharing (as long as there are not profit) and supporting authors are not in direct confrontation. In my humble opinion and personal experience, they are correlated.
"Private sharing" and especially recommendations are also my favorite ways to find worthwhile books.
Recommendations are great once you both have some favorites books in common. Because of that, I always check the Amazon's "Other people also bough" section.
I just wish that jQuery had support for XPath style selectors as well. Chainable XPath would be hella sweet.
We also have a separate Phantom/Casper script which we use to test that our Twitter login flow is working.
looking forward to being able to distribute jobs across multiple machines