Harvesting data from websites using WebKit and PyQt4
rkblog.rk.edu.pl
rkblog.rk.edu.pl
Using WebKit in PyQt4 we can write an app that will collect data of all ads on a webpage and parse the data for marketing guys .. [snip] .. Next stage will be to make application that will open ad URLs saved in DB
NO! Any half-decent ad-network will quickly detect the unusual clicking and will freeze the site owner's account (While still billing the advertiser for the click!) My ad engine will learn your behavior in about 5 clicks and after that, welcome to my de-optimization hell, hope you like PSAs, non-profits and humanitarian causes (until you became too much of a nuisance, then you're a "drop" rule in a proxy filter.)
If you're trying to explore an ad network's inventory for competitive advertiser poaching (you want to bring their advertisers to your site) you need to just save the banner ads and use actual humans to read the ad, google the firm and contact them. Every advertising link you see on line is a 302 redirect and someone is paying for it, usually the advertiser, and if you click on it fraudulently like this, both advertiser and publisher. Besides, don't refresh the same page; the ads to that page are already contextually targeted, and you only see matching assets. Instead, hit multiple pages on multiple sites, and each few times to get the big picture.
There is no bigger danger online than a moron with a "for" loop.
This tutorial shows whats possible with WebKit, which isn't possible with CURL :)
The networks can detect fraudulent clicks easily, it's the site owners who will have their accounts frozen.
http://www.rkblog.rk.edu.pl/w/p/harvesting-data-websites-usi...
http://www.rkblog.rk.edu.pl/w/p/harvesting-data-websites-usi...