Spam vs Mahalo: Matt Cutts Explains the Difference
seobook.com
seobook.com
the catfights landing on HN (and getting voted up) is doubly annoying!
Double standards need to be exposed.
If you do any of these things that Mahalo is doing, you'd get delisted from Google in no time.
That is bad news for anyone who tries to get long-tail organic search traffic. It's good news for sites with great brand names, but terrible news for anyone else.
This kind of frustrates me to the extent that I compete with Demand Media (limited, but real), because there is absolutely no circumstance under which a Demand Media article is a better result for the user than a page on my site if we're in competition, but Demand Media has the economics of it nailed such that they can produce content for some genres at scales I can't possibly keep up without duplicating their core strategy (automated keyword selection -> cheap freelancer -> post without substantial quality control), pushing my related content out of the SERPs and then charging me for the clicks I would otherwise be getting for free.
AFAIK, Google views their relationship as symbiotic, not parasitic.
When you think of how many struggling freelancers use those long-tail guides to build their business ("How to shoot a commercial for a gym," or "How to write brochure copy for life insurance,"), you can see the magnitude of this problem. People who could trade their time for traffic now have to trade their money for traffic. When they're just getting started, money is harder to come by than time. The result: fewer people creating this kind of content, more of them joining organizations that pay for the traffic instead.
All in all, I still think this is primary a Google problem, not Demand's. They can publish anything they want, it's their right protected under free speech. It's Google which should be concerned about quality of their pages.
You are correct about this being Google's problem. These guys exist to exploit an arbitrage opportunity: Google's search algorithm picks them, and the average searcher's which-engine-do-I-choose algorithm picks Google. In the long run, one of these things will stop being true.
Their model is similar to Demand's in that it is UGC + payment, but that's about it. The topics are not generated by an algorithm, for instance.
I was surprised, too.
The result isn't Pulitzer worthy, but its a lot better than you're making it out to be.
Using your example, here's the wikiHow article on "how to make pancakes":
http://www.wikihow.com/Make-Pancakes
Its quite informative. How would you suggest they improve it?
There is value in duplicate content. But there's no value in a search query taking you to a search page, where the site you wanted to land on has to pay Mahalo for adding an extra click between you and what you wanted.
Google would rather make a cent each on a million stolen pages of crap than $10 on an original content, plus the spam "author" is not going to ask anything for the "content".
Every time you post anything about it, you prove you don't deserve to exist. Or you are a spam bot.
Will anyone benefit from having read this article?
Are there really that many people interested in this?
Are you all really up voting this stuff because you're legitimately interested?
If Mahalo (or anyone else with a high PR domain) can outrank everyone else simply by spamming, it's not fair to a nascent startup
These are topic pages that people are working on and THEY DON'T RANK in search engines until they we get the word count to around 300-500 words.
We are the process of NOINDEXING the pages that are below 300 words just to make Aaron happy... we actually had these noindexed before our last version and that got lost in the shuffle of the new launch (really, it did... when you do new code you might leave something out of the old code).
i'm also getting a list of every page under 300 words and having the page managers build them out in 30 days or deleting them.
Anyway, i thank Aaron for busting out chops and making us better!
The claims that we are "scraping" are absurd... we're using google, bing, twitter, etc. apis to do a comprehensive search page.
i dont know everything about SEO, but i don't understand this claim by Aaron. i think he is trying to start trouble for us... and maybe it will work. Thanks pal!
Some people would consider that scraping. How do you define "scraping" and how does this practice not fall under the definition?
- If you don't want those pages indexed in Google then why are you submitting them in an XML sitemap?
- I have already shown examples of the 0 original content pages ranking, so how can you claim that they do not rank?
- You are not scraping directly, you are pulling from 3rd party sites and using it as content on your own site. Which is worse, because there is no way to opt out of it.
- My problem is not just with what you call stub pages, but with most of your pages. When you give people embed code to embed your content in their site you give them an iframe AND a direct link back to you. If you want me to stop highlighting the absurdity of it then perhaps you should hold yourself to the same standards as what you offer others. But you do just the opposite when you embed 3rd party content in your site. You slap a nofollow on the links and embed the content directly into the page (rather than in an iframe).
- Worth noting that every time I mention the above point you end up talking about stub pages or experiments or some other strategy to try to redirect attention. But in reality, what I am talking about is what you do on almost every page of your website.
2. they don't get traffic is my point... we look at any page that gets over 100 page views in a month and we build those pages out. so, even if you find a page that ranks it will not have traffic. if it has traffic it gets built out.
3. we are not scraping, we are using search APIs
4. i dont understand this issue of our widgets (which don't get used to be honest.. it's a failed program)
5. this is simply false... our traffic comes from how to articles, walkthroughs and Q&A. if you want to know what the top 10 pages are they are things like how to play guitar and call of duty walkthrough pages. those things are 3-5k words!
just lay off dude... go troll someone else.
I'm curious! I'm in the content-creation business, and if what you're doing works, I'll either need to radically change what I do or to start copying you.
Ah, so now you admit it was intentional. But good on you for (eventually? hopefully?) fixing it.
- 2. they don't get traffic is my point... we look at any page that gets over 100 page views in a month and we build those pages out. so, even if you find a page that ranks it will not have traffic. if it has traffic it gets built out.
If a person has a quarter million pages that are getting 5 visits each that is still a lot of traffic. Especially when the page has 0 editorial costs.
- 3. we are not scraping, we are using search APIs
The end result is what people would typically call a "scrapper site". It is irrelevant how it is created (if you scrape directly or syndicate from somewhere else that is scraping). The issue is a lack of editorial control (see your page about 13 year old rape) and a lack of citing sources with links.
- 4. i dont understand this issue of our widgets (which don't get used to be honest.. it's a failed program)
Search engines have duplicate content filters. If the content is within the page as HTML (as you do on Mahalo) then you can often outrank the original source for their own content. You bypass this issue and me mentioning it if you only use an iframe to embed the content in your pages. But if you embed it directly into the HTML (as you are doing right now) then of course it is bogus.
- 5. this is simply false... our traffic comes from how to articles, walkthroughs and Q&A. if you want to know what the top 10 pages are they are things like how to play guitar and call of duty walkthrough pages. those things are 3-5k words!
I am not talking about your top 10 pages. I am talking about the bottom 300,000 pages, which in aggregate get far more traffic than the top 10 pages do. :D
- just lay off dude... go troll someone else.
Not trolling at all. Just trying to give you valuable feedback, as you have claimed it to be publicly multiple times (unless you were lying when you stated that) :D
this will all be done in the next 72 hours and then there will be nothing to complain or write about after that Aaron!
Thanks for making us better.
Does that mean that (for the remaining pages on the site)...
a.) the other scraped content which exists on the remaining pages will be put in an iframe (rather than as text on the page)
- OR -
b.) that you will be removing nofollow from the pages you are scraping content from?
Either you trust the content enough that you should link to it directly, or you should put it in an iframe such that search engines don't see it. Either route would likely be more akin to fair use than what you are currently doing (automatically scraping 3rd party content into your pages and using it to rank against the content creators, without permission, and without a way of opting out).
1. Breaking news: http://www.mahalo.com/nhra-fan-killed http://www.mahalo.com/andrew-koenig http://www.mahalo.com/bloom-box
2. How To articles http://www.mahalo.com/how-to-speak-french http://www.mahalo.com/how-to-play-guitar-for-newbies
3. Walkthrough articles with our videos! dozens of them here: http://www.mahalo.com/walkthrough http://www.mahalo.com/call-of-duty-modern-warfare-2-walkthro...
is there something wrong with this pages?
It's been a few weeks: have we seen any evidence of this happening?
Likeso http://tinyurl.com/ydyo3ud
It's not just Aaron Wall that sees something fishy.
But for those of us who are just tired of the whole drama, just change how it's done, or don't do it at all. Adding nofollow and not submitting auto-generated content in Mahalo's sitemap does not seem like a great amount of development work if you really want to change it.
Similarly, Hacker News homepage might be considered spam because it's mostly titles of articles. Until you take into account the value of the votes--which is value that these google guidelines don't address clearly.
I don't think Google's search results or even Hacker News really ever shows up too often on SERPs except for their respective brands... so it doesn't really apply. They are destination websites not focused on search traffic.
For example, links such as http://topsy.com/s/toyota+grand+jury+subpoena are submitted into google to be indexed.
This is a bummer as this just serves to increase the noise to the detriment of getting good results.
I don't think so. It looks like the policy is in place to keep you from having to perform multiple clicks to get where you want. In the common use case (I visit Google.com, type a search, get a result) I do not have to click on anything more than the search result to get where I am going. The difference is in a use case where my search result takes me to a search results page, which means I am only halfway (or less) to my target page after clicking on the Google search result. This doesn't mean, necessarily that all Mahalo content is bad, but that Google shouldn't be indexing Mahalo search results pages, and instead should link directly to the resultant content pages instead (scraped or not.)
Similarly, Hacker News articles should not be the result of a search done on the news title (as HN's link to the original article actually increases the article's score, right?), but should perform better if the matched term is found in the HN-specific content (e.g., the commentary we create on HN only.)
I'm sure that there are (hopefully) edge cases in which a search for the article title returns the HN link instead of to the original article, but that should be the exception to the rule. Of course, if HN were to have an absurdly high page rank and link to a relatively unknown (or new) blog, it would likely lead to a result on here, instead of the other way 'round.