Trolling the search engines
blog.dam.io
blog.dam.io
The theory behind it is that if your site has a lot of relevant "targets" then it must be more important than a site that has only a few targets. (Consider wikipedia as the poster child for this)
When people write naive web spiders in an attempt to create their own crawler, sites like this 'trap' them in an infinite web of apparently unique links. Always a good idea to stop after a few hundred and kick it back to a human to see what is up with that :-)
This has certainly been done before: http://en.wikipedia.org/wiki/Spider_trap
I had to kill the experiment (no more new "pages" crawled) because of the CPU load and bandwidth costs, even throttling robots
Circa '98 IIRC, there was a module available for a webserver I worked with which generated a page with number of bogus email addresses per page, and a number of random urls per paged that when followed generated yet another page of bogus email addresses and links.
You hid the link somewhere on a legitimate page, added the base path as an exclude in robots.txt, and any mail harvesting spam-bots would get sucked in.
The idea may well have been around longer still.
This is the core of why Google uses things like links as a ranking signal. If they based it entirely on the content of the site, they're easily duped. While it's somewhat trivial to manipulate your backlink profile to increase your rankings, it's rather hard to fake links from known high quality sites like the New York Times for instance. So they can "trust" links more than they can trust the content of the site they're crawling.
So even though this site has a large amount of content indexed, the changes of it ranking for anything more than gibberish are so low that it doesn't even matter.
Though it didn't last very long, Google penalized the websites after couple months. I did it for testing, but it could be an effective strategy for spammers who could rinse, repeat and scale.
There are many websites like that still ranking and generating traffic; some of them have Alexa rank below 1000.
http://inf.demos.dam.io/sdfsdfs
I have yet to find a pattern, but I guess the author did some sort of hashing to the query string and use the hashcode as the seed to generates random text. For some given strings, the backend fail has to do with the use of the hashcode probably. (e.g. use [some_variable_or_constant_here]/[hoshcode - CONSTANT], when hoshcode == CONSTANT, the backend code does not contain the exception handling code.) Just a guess.
If you want to impress us, rank a few 100k pages for competitive terms with garbage content.
EDIT: Installed AWStats, just wait an hour: http://stats.demos.dam.io/
EDIT2: Can't manage to make AWStats to parse the old logs...
Btw, does anyone know who this ip(216.151.137.36) belongs to? All I get is spam reports when I google it and doesnt seem to belong to any (major) spider. http://stopforumspam.com/ipcheck/216.151.137.36
Also on another question, how does google deal with web apps that have unlimited pages(dynamic urls based on get and so)? As in, how does it say "this is legit" and this site here is not. Surely backlinks are one thing, but these can be "faked", too?
FWIW, I see this ip hitting my spider trap too. 14 requests in 30 seconds on March 14th, and 21 requests in 45 seconds on Feb 20th.
am i wrong in thinking that might actually move SERPs?