Too Much SEO? Google’s Working On An “Over-Optimization” Penalty For That
searchengineland.com
searchengineland.com
Too many sites are ranking by stuffing keywords and buying links.
Next, Google need to find a way to rank a page based on it's role in a broader user experience - rather than being measured in isolation.
Sites (especially in retail) are forced to optimise a page in a specific way to get indexed successfully. The result is everyone putting too much content on a single page - detracting from usability and creating homogeneity for users - like cars all optimised in wind tunnels.
By buying links (undetected) you can manipulate the quality and relevancy of your website in Google's eyes and cause it to rank better than it should.
I think algorithmically they can only discount links from bad neighborhoods, etc... but if they catch you, they will manually penalize you.
I think we are a ways away from social signals being a "big factor." While google does claim that they look at social signals, I think from a practical standpoint the only value they provide is raising a red flag when someone manipulates social signals like buying likes and plus ones. The same way google looks at a link profile of a site, they can catch a bad social profile that seems unnatural.
In reality, links are almost never seen as a negative. Which is really a shame, because it makes it harder for good-guy SEO's to legitimately build links when you can just buy 10,000 on Fiverr and not worry about it. And most black-hat SEOs buy links from many players, so even if some are discounted they won't usually still come out on top.
Google is getting better at it though. Recently they've been cracking down on blog networks. This is purely conjecture, but I do believe that links are on their way out as the primary signal.
That's the whole point of the Google backlinks - get as many links back to your site as you can, and you'll get good SEO.
And if you do this, and Google detects it, they nail you.
I'd never do it; it's just an interesting question.
And the implications of that make me shiver a little.
(Disclosure: I have no idea how much it would cost. But my hypothesis is that it would be the same as if the website owner had done it.)
People who have a high PR website may sell a link (usually in the footer, on all pages) to websites in a similar niche to help them rank better. However, for some people its not blackhat enough. Spammers nowadays tend to use automated software like the following:
-xrumer - Developed by a few russian people and primarily used by russian pharmacies for fraud, this is the most powerful "spammer" too out there. You can literally spend $1k and get 1million backlinks within a matter of days and have your site in the #1 position for whatever keyword you're aiming for.
-scrapebox - Similar to xrumer, but while xrumer targets forums (vbulletin etc), scrapebox specializes in harvesting large lists of wordpress blogs (throught google + proxies/botnets) and then posts comments on them (usually all those generic comments you see on wordpress blogs such as "nice post" "good advice, bookmarked" etc).
I could literally go on for hours and hours on how people are exploiting search engines, I think that its prime time for Google to develop an AI for this rather then just relying on algorithms. There needs to be a better PageRank alternative that needs to be implemented.
>benefits
In the most basic terms, for spammers, more links = better search rankings = more traffic = more $$$
Besides, if one really good one does exist, many/most SEO people would be out of a job; white or black hats.
Should the human reviewer check every link by hand?
Think codecademy.com now first for "learn to code". I mean, is it best place to learn how to code? At the moment, I'd say it's not. Is it the third result "codeyear.com" (run by codecademy as well)? I'm sure there are better sites that could take that spot and provide more value to the user than that landing page pointing to codecademy (Nothing personal against the codecademy guys of course, big kudos for their venture so far and best wishes for the future).
The problem with today's algorithm is that it can't really tell which site is really good and which one is just popular, mainstream, consolidated, with a high PR. Think lifehacker.com second for "learn to code" (on my side of the world, on google.com and chrome incognito mode on).
Right now if every single newspaper of the world cover your site about how to build a homemade nuclear device and actually write something about the matter in the article (how it would be natural to be), use as anchor text "how to build an atomic bomb" and link to your site, your site will skyrocket to the top for that kw (and many others similar). Even if it's just a school project or a joke that made it to the news.
Where's the quality check?
Anyway I'm sure that smart people are currently working on these details and that a comment reply can't really do justice to the complexity of the issue.
Sorry bro that ain't true anymore. Xrumer brutal blasts used to rule years ago, but I can assure you that just having a million of crappy profile/forum post links popping up at the same moment nowadays isn't going to help your site ranking as it used to be. It's just not as powerful as it was before the whole xrumer was translated and sold to non-russian speaking folks.
Xrumer is still being deployed by the "big dogs" but not to link your main, "money site" directly. More like to link to other pages with links pointing to your site (or to pages with links to other pages with links to other pages with link to other pages that eventually link to your money site). Never directly, too risky (even though it's debated how risky it is, but still why taking a chance to find it out).
Buying links today means (more like) purchasing services like seolinkvine.com or buildmyrank.com where you buy posts (about the kw you want to target and have on them a contextual link) on network of sites with an high home page Page Rank.
Nonetheless Google is catching up with these networks, recently deindexing a lot of their sites and making their users, who spent lots of money for those links, feel really, really, really bad.
Pros in tough industries had to adapt to Google's algo evolution: they now build their own networks, they put out decent content on them, they provide a good user experience, they tend not to over-optimize, they host them on different A-class IPs, they have their domains registered with different names/addresses/etc, they build multiple tiers of links to their backlinks with custom developed softwares (or Zennoposter) with private proxies and their 24/7 running servers and so on.
It's not a cheap process nor a quick one.
As for now and long-term rankings, spamming isn't just enough.
To be on the safe side you really need to follow Matt Cutt's evergreen advice: build a kick-ass site that even a human reviewer working for Google would love to see on the first spots.
Knowing how easy it was to manipulate google results (and to some extent, it still is) imo that's the right direction from the perspective of a everyday user who's searching for something that's not clicking on Adsense ads or buying something.
Just to be clear, there are 40 some factors that effect the value of a link, and that is just on a generic level, these factors carry different weight in different verticals, and based on other signals. http://wiep.net/link-value-factors/
I think we need to wait and see how google rolls out this new penalty to see how it plays out.
Just to explain, there are over 200 publicized ranking factors and I would bet that each one carries a different weight based on all sorts of signals google sees about your sites as a whole, your competitors sites... So, a spammy old strategy like the ones mentioned above are hardly a strategy for ranking, they are just working in those instances because of broader signals at play...
A simple example of this would be... Most SEO's agree that page titles are one of the main ranking factors as far as relevancy goes. But if google sees half your titles are duplicates, they might discount all your titles across your site because they determined it to be a bad signal in your specific instance...
I took a beating on the SEO blogs for calling SEO a bug a few months ago, but I'm glad the rest of the world is finally realizing that I'm right :)
Ultimately, Google is trying to rank you highly for providing the best content; you shouldn't be spending your time trying to figure out how to game Google by making superficial changes to the presentation of your content. Want to rank better? Write better!
The whole problem is rooted in the fact that Google is a leaky abstraction. It tries to be an omniscient Sherpa, guiding the wary Web traveler with its infinite understanding of the Web and the individual user's needs. The reality, though, is that Google is actually just a computer program. So there is a gap between the user's mental model of Google and what Google actually does, and it's this gap that SEO exploits for its own profit.
An infinite amount of exploitation would mean that Google would just return results randomly, and so it makes a lot of sense to detect signs of SEO and penalize the behavior before it further broadens the perfection/reality gap. Gaming the system is currently profitable, since the worst thing that can happen to you is nothing, but the best thing is that you get more traffic. A penalty aligns the risk/reward spectrum to favor "write better content" rather than with "spam a bunch of wikis".
Obviously Google doesn't think their search is good enough, and I would agree -- piling more "inputs" and arbitrary branches into a ranking algorithm, however, is no solution. This will only devolve into an endless game of cat and mouse until a new search engine comes along and does to Google what Google did to Yahoo.
The real world doesn't always work like that. I do a lot of work for an eCommerce site that sells wholesale. Their customers have little to no interest in reading text, all they want are pictures. Which means, the catalogue pages, which are optimised for actual human visitors and not Google robots, contain little to no text, only pictures. As a result, the catalogue pages (most important part of the site) do not rank anywhere with Google, and never will. Google is unable to handle websites that are - quite correctly - all about the pictures. Competitor sites that outrank this particular site design their catalogue pages for Google, not for humans, and rank well because of it.
I'd rather focus my energy on driving traffic in other ways, such as traditional media, list building, PR - and focusing on creating great content.
I just do the bare minimum stuff when it comes to SEO (title, h1 tags, some anchor text here and there). Other than that, I don't pay attention to SEO in the actual content of the pages.
Create content & do minimum SEO.
Eventually, other sites will link to your content & Google will figure it out.
If you try to stay on top of the latest SEO tactics, you end up hurting yourself in the long-term. For example, awhile back it was recommended to create huge directories of conten t (something like: Travel deals in Ohio, Travel deals in NY, etc, etc). With Panda update, you are now penalized if you have too much similar content. Why deal with all that headache, just create good content.
But telling someone to "just write good content" is like telling a programmer "just create a good product". The world doesn't work like that. You need to hustle, shout and say Look at how awesome this site/app/article is! Link to me! I don't like it that much, but that's how the game is, and if you don't play it accordingly, you'll end up with awesome content and no visitors.
Traffic to (and engagement on) a site seems to be a much better signal, for the many sites that use Google Analytics, at least. But not everyone uses GA though...
Backpedalling now, PageRank is about weighted links. Link farms should have no weight to give, and links from the user generated parts of pages (which are easy to detect on most blog platforms) should be easy enough to weight fairly too. So it is hard to see why Googlebot gets fooled by any blackhat link techniques that Google humans are aware of, except maybe the case of selling links (influence peddling), but one can argue that is a social /legal issue, not a technical problem.
Taking a look at the site the domain registrar wasn't doing a proper nameserver/cname/permanent redirect or the like, instead putting the shop host into an old-skool frame. Not only that but for some bizarre reason they created a subdomain and were framing that too.
So a frame inside a frame inside a frame. It's genuine hard faults like there that's the reason SEO will exist for the foreseeable future.
They still rely on this dumb word-based approach to document retrieval. Example, if my page is about how to manage your time, and it's a really fantastic resource about that, but I don't actually mention the phrase "how to manage your time" anywhere, I won't rank for people searching for that phrase. I should, but I won't.
So I have to write my content for two audiences - humans and Google.
I don't want to do this. I'd be much happier just building great content for my human readers, and if I happen not to mention the exact keyword phrase a searcher might use, it doesn't matter, Google still knows it should put my site at the top for that phrase, because it understands that my site is about that phrase, even though I don't mention it exactly.
But Google just isn't good enough yet. Someone can set up a page which has a bunch of headers and URLs and variations of the keywords and beat me to #1, even if their content is utter rubbish for human consumption.
That's why SEO still exists. It's symptomatic of a bug in Google.
This penalty will help, hopefully, but we're still going to be in a state where we have to compromise our content to serve two masters.
That's really what I was getting at. Stripped right down, Google is still just viewing documents as a bag of words[1]. I mean, they have pagerank and they will apply more weight to words in headings, and they have 6-gram indexes and synonyms and all that clever stuff, but at it's core it's still lexically centered not semantically centered.
[1] Further reading: http://en.wikipedia.org/wiki/Bag_of_words_model
Google could stop bolding the keyword matches in the SERPs, but that would be a disaster because they are useful for picking the best result.
My point is that this is not purely a Google bug, but also a bug in human nature. People like exact keyword matches to their queries.
I'd love to see a source to back that up. The only data I've seen on this is various eyetracking and/or click tracking studies which suggest higher up on SERP=more clicks.
This is absolutely correct!!!
I watched the video they released on improving search quality yesterday (http://insidesearch.blogspot.com/2012/03/video-search-qualit...) and was initially impressed that they gave so much thought to slightly improving a small subset of 0.1% of queries. That meeting must have had 40 people attending!
But afterwards it left me with the feeling that Google is becoming so big and entrenched in old ways of doing things, that they may not be focusing enough on the next big improvement in search. Penalties and heuristics can only go so far -- eventually they'll need something approaching AI -- and if any company can do it it's Google.
If you were thinking more along the lines of: "you ask a question and it understands your question and answers it like a human would", then I can tell you that this would require a scientific breakthrough first. That is not really something that you can plan for. Also Google mainly excels in engineering, less so in research. I actually think that Yahoo and Microsoft have stronger research divisions.
I'd also like to note that for people creating content on the web, especially programming-type people who deal in logic and certainty, the SEO system resembles something like black magic. It's pretty clear that unless you understand SEO and apply it, you're never going to be seen no matter how good your content is. Now it appears that if you understand it too well that's also a bad thing.
The goal here is to let the users themselves inform the search engine as to what content is good -- hence the plus-ones, social search and all of that stuff. But all of this is still indirect evidence. Unless you could plug a computer inside the head of a person and watch their every thought, the only real data you have for input is server logs, click-throughs, and all kinds of other things that computers do, not people.
I just don't see this being solved any time soon. But I do see it getting so complex and unwieldy that it continues to frustrate searchers and content producers alike. Meanwhile the bad actors will continue to have a heyday.
Wish I could be more optimistic about it.
EDIT: small clarification
Admittedly, that is easy for me to say because I don't need high traffic at my site. I need my web site to have a few high value visitors: people who want to work with me or communicate because we are into the same technology. The way I "optimize" my site is writing about what most interests me, and that attracts people with my interests. Seems pretty straight forward to me.
Regardless, you are taking an overly simplistic view of SEO...
SEO works best with quality content backing it up.
As an aside, the first over-optimization that I would target if I were google are keyword-based domain names. Keywords in the domain name are given WAY too much weight; how often do you find that the top three results for the search "XXXX YYY ZZZZZ" to be very shallow but keywords-rich websites with the domain names www.XXXXYYYZZZZZ.com, www.XXXXYYYZZZZZ.org, www.XXXXYYYZZZZZ.net?
EDIT: But I completely agree. There are far too many SEO's out there for this to beat everything
Here is my search process
If it is from ehow ignore
If it has my search term in the URL ignore
If it is from about.com ignore
Do you have any examples of search terms that don't return useful results until page two or three?
From what I've read over the years, basic SEO mainly boils down to :
- Have good, relevant content
- Choose your page titles and URLs carefully
- Get lots of links to your pages , from as many different domains as possible, and try to get them from high Pagerank domains
- Where possible try to get anchor text relevant to the terms you want to rank for - but don't go overboard with a large percentage being the same phrase, as it appears artificial
- An older site can benefit you, as will exact match domains for the main TLDs
Except for site age, they are all easily changeable by the SEO.
Maybe Pagerank is too easily gameable, and what is old will become new again, and some of the approaches tried in the 90s and replaced by Pagerank will return with a new twist.
This is done mostly by identifying related / tangential / complementary / supplementary keyword terms to the resource linked to. Non whitehat SEOers are wising up to the idea of a site targetting niches, and that considerably widens the list of appropriate keywords.
By tackling the long tail keywords en masse - these are less competitive, easier to rank for, and correlated better to buying intent; gradually they make significant inroads to ranking strongly for the big-head keyword in a manner that looks more natural.
And, apart from that, Google totally sucks on languages that are different than English. They think of a new thing for the English based search and then they'll just propagate it everywhere with no substantial modifications whatsoever.
I've got many sites under me, all writing original content. They're not gonna win the Pulitzer, of course, but at least they're edited, accurately followed by teams of human authors that won't even copy PR as a measure to prevent non original contents to appear on the sites. Although all of these things, we got seriously pandalized.
The results where we once used to stay atop are filled in many occasions by scraper-sites who steal our contents and rank 2 or 3, sometimes up to 5, positions above us on the SERP. I can't even count the times I filled the anti-scraping form anymore. That's not enough, because many times popular sites rank higher than us just because they're pretty popular, although their contents are horribly written, short, totally uninformative.
That happens all the time. You know where we still go pretty strong? On Google News, where the human intervention sometimes really applies.
Google should see what's going on everywhere and their insistence on having matt cutts as its only public voice on this huge issue is becoming really frustrating.
He ends up looking like a fake good guy, perpetrating the hypocrisy of a corporation too big and too convinced they have the ability to solve all the hyper-complex search problems, laden with human generated unpredictability and the natural human tension towards deception, just with pretty algorithms.
Firstly, for ecommerce sites etc due to the lack of content - especially unique content (because there are only so many ways you can say something is X length) these sites are forced to SEO themselves. This is even more apparent as Google are placing their “Google Product Search” within the results as well. Sure, you could add yourself to Google Product Search but that’s not the point - Google should be adding ecommerce sites etc in their automatically to make it a level playing field. Sure, you can argue that it is hard to do and that might be the case but, there is proof out there in the marketplace to highlight that this is possible - look at what TheFind etc are doing in this space. Additionally, I think there is a lot more that can be done in the Shopping Search space as well as other areas of search which I will cover below.
Having covered those, I will also highlight another problem which is hindering Google = their search engine is based around the Pagerank algorithm which despite evolving is actually hurting Google in trying to solve the problem of SEO.
I believe this is the case because of PageRank and the general Google Search algorithms which are in place – their search at its core is based on 'citations' like those in academic papers etc (yes it has evolved over time) but it is still even loosely based and developed upon on this ranking system. Hence there is the problem of paid links - although you can report them [1] this reporting method actually isn't effective and doesn't work. Additionally, there are tons of ways to get links really easily which appear natural, won’t appear overly seo’d and are extremely easy to get and game as well.
Google seems to be taking their search into the, Semantic & Social Search approaches which they will probably solve some of these aspects but, everyone is working out (and many have already worked out) how to game social search although, it is a step in at least a decent direction.
Currently, Google just provides links and documents etc to the ones which it believes is the most useful based on their Pagerank/keyword approach and even if it can improve search to a much greater degree the issue is also related to Adsense/Adwords.
This is because; these two print money for Google. For instance, Adwords really is used for Google Search (yes it does feed into Adsense as well but Advertisers really want their ads on Google Search Results). The last statistic I heard about Google Search, is around 2007 when Google was making $0.12 on average per search – you could probably calculate easily how much Google is making now by looking at their income and dividing it by publicly available search numbers but, I’m going to say it’s more than that because in 2004, they were making $0.09 per search. This is actually an issue for the user because, I believe Google really just want their ad results to be perfect – they don’t really want you to have to click on the result because then Google isn’t making any money – as long as their results are good enough so then they’re Ok with that If you don’t believe me, take a look at [2] that’s a whole lot of ads above the fold. Oh, and if you try and do that be prepared to be punished [3].
Now Adsense, sure it doesn’t make that much of Google’s income but it still is highly profitable – if it wasn’t they wouldn’t be doing it. If Google, really wants to fix search then they need to review every single Adsense site and I mean every single Adsense site since Google only reviews the site which you apply with. This leads of a huge problem as MFA’s aka. Made For Adsense slip through the net afterwards and if you look at their sites, they’re generally have poor content etc and are optimised for the search engine and the user to leave either by clicking the back button. Sure, Google will lose some Adsense income on their balance sheet but if they want to fix search this is also a good place to start.
Sure they're trying to solve this issue and get to grips with it but I don't think they will anywhere in the near future. You can argue this fix they have announced will solve some SEO problems but, as I said in my opening sentence I believe this is the wrong approach and some of the above reasons highlight several issues already but there are tons more. In fact, I actually hope you disagree with me because, not only do I value your opinion but am interested in the other aspects you may add to it.
[1] https://www.google.com/webmasters/tools/paidlinks?pli=1
[2] http://searchengineland.com/figz/wp-content/seloads/2012/01/...
[3] http://searchengineland.com/too-many-ads-above-the-fold-now-...
Google doesn't really know what you're searching for, it doesn't understand the context which is why you see shopping results when you're just looking for information. Yahoo actually tried to solve this with a slider with an option for more shopping/info results but it didn't catch on as users don't want to click anything, they just want to type words in to the box and get the perfect answer - be it an ad or an actual result. If you make a user do anything else then you've already lost the game.
This is another reason why I don't think Google will fix this SEO problem it has any time soon, and I definitely don't think this current 'fix' will solve it either
The deeper problem of the semantic web is that it's meta data, which tends not to be visible to visitors. And because of that, it's easy to forget / overlook / incomplete, and nefariously allows SEO over-optimisation invisibly that non-whitehat SEOers do in full view of the human audience.
I wish it wasn't true, but Cory Doctorow's portmanteau of metacrap is unfortunately still accurate ( http://www.well.com/~doctorow/metacrap.htm ). We have made several small metadata improvements over the years, but not much progress in substantially overturning Doctorow's original concerns.
Human beings are not great sources of accurate and up-to-date metadata. So we rely on scripts and services to fill in the blanks that we don't. (publish times being recorded in WordPress, for example.)
Converting human content into metadata ends up being an automated human-hands-off affair which means that those automative tools need to be able to parse and extract information / meaning from human prose. Very much what Google has been doing since inception.
It's flawed because it relies on automated interpretation of prose. But there isn't a viable non-flawed method of getting the same information without imposing academic-like constraints to the Web.
Mind telling me how a consumer web startup is supposed to get others who are interested in X to visit his site about X? This is the basis of SEO: matching searchers to pages. It doesn't happen magically just by wishful thinking, and writing good content unfortunately.
But because that's the case, and you got a quality product, you're saying to sit tight, don't get out of the building, don't try to sell... and hope people will eventually choose you?
So how exactly is a consumer startup going to get people who want to know X to visit your site about X?
Also, not everyone understands what good content is, or how to properly markup their websites.
The field of SEO is quickly (or slowly depending on who you ask) being transformed into a much more holistic discipline where content creation, usability, conversion optimization, pr and marketing are falling into SEO's job duties. It's much more than page markup and links.
People can say "create quality content, don't worry, you can always file reconsideration request if it's an accident" don't realize how Kafka-like Google really is. If an algorithm penalizes your site by accident, and your traffic drops by 99%.. good luck trying to talk to an actual human being working in Google to get your problem fixed!
- Less likely to find results in forums
- Harsher on results from “content mills” (even after Panda)
- Less likely to give you ranking for a keyword if is not found verbatim in the page
- Also provides way fewer results for phone number queries
Creating good, worthwhile content is a key part of SEO. If you write, or create, good, worthwhile content, you are taking part in SEO. If you take your time to write a headline that best reflects your article, and mark it up correctly, that's a part SEO. If your site is using easy to read names, that's SEO. If people link to your content and write about it, that's a part SEO. If you make the description below your page title on Google's search results actually worthwhile then some random piece, that's a part SEO.
People who "over optimize" on a large scale are probably the same people making money off of well ranked SERP's, penalizing those people pushes them to buy adwords.
They are in essence pre-qualified customers for adwords since they have already demonstrated a willingness to invest real time/money in ranking. Sure this will just trigger yet another race to optimize optimization, but until things are refigured out more money will be dumped into adwords.
Thought 2: This SERP "Market Volatility" is a great way for Google and SEO's to make a little more money.
This is also a good way of reminding some people how much their revenue is dependent on Google, when you see an overnight dive in revenue it captures the attention/mind share of higher ups who have maybe been taking their well oiled SERP machine for granted.
It's like a one night only "Google Dance" reunion tour, and gets people obsessing over them again.
What will be left will be personal bloggers, charities, and hobbyist sites. Be careful what you wish for.
To clarify for your benefit, I wasn't complaining about Google's optimization or lack-thereof (spammy vs not spammy). I think it's hilarious, watching the endless dance that Google is going through because their fundamental approach to search is broken and they're trying to drag that broken approach into the future. It's like watching Microsoft with each iteration of Windows & Office, trying to figure out how they can cheat death and drag 1980s software into the future.
On a business level, I don't care about Google's survival or their optimizations. My product doesn't benefit from SEO, nor from their search engine. I'm indifferent to them. If they make a good product, great; if they don't, someone else will eat their lunch eventually.
I wonder what happens 20 years from now on, Google can't win over SEO in the long term. There are just too many people on the other side. As long as they don't develop some AI the results will get worse. I use more often "site:" to get what i want.