“A spambot barfed its post recipe on my blog”
twitter.com
twitter.com
https://www.google.com/?gws_rd=ssl#q=%22time+to+make+some+pl...
Looking at this link (http://prophesyagain.org/radio/#comment-98), it looks like this spam was left for a URL that redirects to http://www.itunescoms.com/, a fake looking iTunes knockoff that probably drops all kinds of nasty adware/malware on your PC.
For those unindoctrinated in SEO spam, the trick is to get past the spam filter, then leave a link in the name/username. You generate hundreds of thousands of backlinks to one page, which Google considers "votes" for your website. You can either send links directly to your "money site" or you can send them somewhere else and do a 301 redirect, passing the link juice.
Usually a spammer starts using something like Scraperbox and a lot of proxies to find thousands of blogs or forums with open comment fields, then use a script like the OP plugged into something like XRumer (http://en.wikipedia.org/wiki/XRumer). The syntax of the OP is called spintax, and it basically chooses a random word inside of the {}s, creating an infinite number of comments. You find a few hundred thousand open comment fields to post in, ride a little wave of SEO boost until Google finds you and kills your site, rinse and repeat.
Nowadays most spam syntax as simple as the OP will be caught in filters or penalized by Google, but most spammers are actually really bad at what they do (as you can see by them forgetting a bracket and dumping the entire spintax).
Tip: The result count estimate on the first result page is often off by several orders of magnitude. Often it's possible to navigate to the end of the results pretty quickly, and the number drops to something like 469:
https://www.google.com/search?q="time+to+make+some+plans+for...
1) grab index for each word in the query 2) grab first x results from each index 3) cross these to filter down to actual matches
I believe step 3 is as expensive as rendering the results AND nobody would like to wait for rendering 600k results before getting some of it. This they stick with some simple estimate, like max(size of found indices), or average or whatnot.
I did a search for "hackerspace" (no quotes) and it claims "About 588,000 results." Paging through the results, I eventually got "In order to show you the most relevant results, we have omitted some entries very similar to the 390 already displayed." Even with omitted entries included, there only seems to be 832 results.
Try http://www.google.com/custom?q=hackerspace&num=100&safe=off&... it will show you that they actually found those thousands of results, but only if you search with an older algorithm.
> Sorry, Google does not serve more than 1000 results for any query. (You asked for results starting from 9000.)
Say, for example, that their Search cluster has 10000 nodes. And suppose that the query returns 60 results from the first node itself; so it could multiply 60*10,000 and claim there may be 600,000 results. But when it is asked to actually go and fetch the results, for various reasons, it may not get to that figure; most of the nodes may just shrug and say "we got nothin'".
>If you like, you can repeat the search with the omitted results included.
https://www.google.com/search?q=%22time+to+make+some+plans+f...
indexing / spidering pages that do contain a dofollow
google doesn't completely ignore nofollow
So if they manage to put a link on your blog, the next step is to make spam links to Your site.
Google needs to scrap the pagerank and come up with another metric. Lately I've seen a drastic decrease in quality where 9 out of 10 pages in the top-10 is just garbage.
Even if I copy a whole sentence from a site with low pagerank, it will be hidden deep in like page 4 o 5 in the google search result.
My bigger gripe is how awful it is to get rid of bad links, either as a result of your own past discretions (or past discretions you've inherited) or from negative SEO.
The editors had started clamping down and banning users for spamming the forum so I suspect some agrived spammer reported us to google out of spite.
I complained and explained it was only a matter of time before the whole thing fell through and got hit by an update or blacklist, but he didn't listen. I did at least redesign the main websites to be html5 with good metadata to mitigate it.
This is what happens when C-levels think it's still the dot-com era.
[1]http://en.wikipedia.org/wiki/Article_spinning [2]http://thebestspinner.com/
Imagine a 100$ tool that might return 0 to 200$, the tool author sells the dream, and collects a more reliable income.
You may see software like TheBestSpinner, XRumer and another one I can't remember the name of, basically a Wordpress comment spammer cannon that sell thousands of copies, more than that person would make using these tools themselves.
Of course, your $100 price might be a figure pulled out of thin air, not an accurate one.
So, either your cheap tool works and spam gets worse or it fails and people just use more expensive options.
PS: Spam is a billion dollar industry and plenty of people with fairly deep pockets.
That said, these blackhat forums tend to be full of people with little money. Many users are from places like India and Pakistan, trying whatever they can to eke out an existence. Since almost none of the users on those forums have money to buy the tools (even $100 is a very high price point for products on these forums), writing tools for them is a fool's errand.
I don't know how either one works well, but I saw some similarities, and with my limited knowledge of how that name generator format works I would know either more or less about how article spinning works.
Waitress: Well, there's {egg|bacon|sausage}, {egg and bacon|egg and spam}, egg {bacon|sausage} and spam, {spam|bacon} {spam| spam spam} sausage and spam, {spam|egg} spam spam {bacon and spam|sausage and spam}; spam {bacon|sausage|spam} {spam|spam bacon|spam tomato bacon} and spam
Subject was: "Dear java.lang.NullPointerException, you have been selected to receive..."
The company I'm working for has alarms which monitor log files for "FATAL" and the like, some of them also contain user input, so naming your kid something like this is another good way to annoy some ops people :D
There was also an attempted attack on the Swedish voting system via a handwritten SQL injection attack, but that was unsuccessful[2].
[1] https://drwetter.eu/amazon/storedXSS-vuln.at.amazon.html
[2] http://alicebobandmallory.com/articles/2010/09/23/did-little...
I'll right away clutch your rss feed as I can not to find your email subscription hyperlink or e-newsletter service. Do you've any? Kindly permit me recognize in order that I may subscribe. Thanks.
re.compile(r'\{([^}]+)\}').sub(lambda x: random.choice(x.group(1).split('|')), tmpl)All the best,
But, if one is dead-set on using pastebin.com, the least they could do is post the raw link so users can escape some of the horrible.
When Google started to use machine learning to filter out machine generated content this out this method quickly lost effectiveness. However, it is more effective on less sophisticated search engines so using this will show some benefit.
One of the more current methods is markov chain based article generation with provided keywords. And perhaps more importantly paying people to write content.
[0] With internet proliferation this has dropped substantially.
[1] http://www.hanselman.com/blog/ExposedABlogCommentSpammersSou...
Example: Name: "Red T-Shirts" URL (optional): "http://red-tshirt-site.com/"
Then they'll perform better for that keyword (maybe--google eventually penalizes this).
Basically they would spam these out to not look overly spammy, and have their website and name listed as their website they are trying to promote as people use to believe that having millions of shitty backlinks is good SEO.
What Google should do is make the links have no effect at all, thus preventing this abuse.
(Although Google denies it's a problem.)
Thank you for allowing me to talk myself into a circle. :-)
- Text that does not clearly market anything or provide a meaningful backlink may be part of a larger link network, or just simply testing to generate a list of vulnerable blogs. Those lists can be sold to others or used for your own purposes.
- Similar to some recent spam bots hitting Google Analytics referrer results, these people may not care about organic rankings. They might be marketing to the blog owner who sees the comment and investigates the username/link. If I'm spamming for SEO services, and you are a blog owner looking to increase traffic for example, the attempt may not be about ranking for SEO terms (a losing battle), but simply getting you to look at my site.
http://craphound.com/overclocked/2014/01/14/when-sysadmins-r...
So what's the payoff here? Based on the text it looks like IE malware, but it could be site visits or back-links or something else.
http://davidsd.org/2009/01/the-real-theorem-generator-a-cont...
(I once ported it to Python and wrote a parser for a more-convenient syntax like the OP's, then never released it partly because I'd hate to see spammers using it.)
whoops