How to scrape data from sites you can't log into
ssscripting.wordpress.com
ssscripting.wordpress.com
Or say the title should be changed to "How to scrape your data from your sites that require login"
For anyone who wants to scrape sites that require login I'd recomend Python with Twill. That lets you do the whole thing with ease.
If you needed Javascript you could use one of the Firefox scripting bridges (Selenium or MozRepl).
I think most people are reacting to the title... I for example thought it was a security posting.
wget --save-cookies cookies.txt ...
wget --load-cookies cookies.txt ...
And Java? Why the hell would anyone write a script in a compiled language like Java? Desperate for that 2ms time saving between 10 second waits for the pages to come down, eh? And any for-real scraping script would have a time delay built in anyway.
The guy doesn't know what he's talking about.
Well, good luck to you, and the more script kiddies you confuse the better I guess, but there are seriously much better ways to do this. Go look at Ruby Mechanize (I think it's also available for python); coming from Java you will be blown away by just how easy this kind of thing is. How do you think we all test? ; )
Update: Oh I see you know Mechanize from another article. So why not just use that ... you do know it can do all that logging in stuff for you, right?
Anyway always good to see everyone chime in with their opinion so thanks for the conversation starter.
BTW, is anyone else nervous about the day the teenage h4xx0rs discover how easy this kind of thing is these days ..
And if it did, how long until a defense is made...
I think once a user is logged into your site, you'd have an extremely hard time defeating this sort of behavior without degrading the quality of user interaction or treading on legitimate use of your site. You may be able to defeat egregious abuses, such as scraping entire photo galleries in seconds, but even that can be defeated by a script that randomizes requests/times between requests.
PS: I recommend all aspiring coders to telent to www.google.com at least once just to feel the magic.
There are frameworks for doing that already, check out Selenium: http://seleniumhq.org/
I'm pretty sure it's illegal in a couple of ways. </bitter>
But yes, the magic is there.
Imagine a web site which terms of service state you cannot use software to circumvent ads. Or where part of the security is done client-side (stupid, yes, but not impossible). Skipping the browser breaches at least the terms of service, and may be constructed as hacking. I think even Google discourages automated searching and prefers you use its api, which (at least some years ago) wasn't free for commercial use. I may be wrong in this particular case, but the important point is you may want to check the specific TOS before skipping the browser.