The most mysterious Google ranking ever...
jamespanderson.tumblr.com
jamespanderson.tumblr.com
STEP 1: Accessing CakeCentral.com returns a 404 "Not Found" HTTP Code when requested:
1. Go to http://www.rexswain.com/httpview.html and enter in http://cakecentral.com/
2. Take a look at the response codes, see the 404
STEP 2: Previously, inexplicably, _actual_ error pages on CakeCentral.com such as: http://www.google.com/search?q=site:www.cakecentral.com%2Fca... returned 302 redirects to Beerby.com
STEP 3: Beerby.com uses a "Soft" error page, meaning if you type in a URL like: http://www.beerby.com/adfadi you get a 302 TEMPORARY redirect to a 200 OK page.
> Many months ago, if you saw someresult.com/search2.php?url=mydomain.com, that would sometimes have content from mydomain. That could happen when the someresult.com url was a 302 redirect to mydomain.com and we decided to show a result from someresult.com. Since then, we’ve changed our heuristics to make showing the source url for 302 redirects much more rare. We are moving to a framework for handling redirects in which we will almost always show the destination url.
http://www.mattcutts.com/blog/seo-advice-url-canonicalizatio...
I thought that the "+" just disabled the spelling/synonym/etc... alterations, while apparently in this query it does some kind of post-filtering for exact matches... (since both words do not appear in the page)
We needed to have a CGI script handle all hits under a certain location, and for various reasons mod_rewrite wasn't an option. So I put something like this in an .htaccess:
ErrorDocument 404 /path/to/script.cgi
I didn't realize until later I needed to explicitly set "Status: 200" in the script's headers. As far as browsers were concerned, everything worked, even IE, since the "error message" (our page content) was long enough to not trigger its built-in error message.
wget http://cakecentral.com/ --2011-01-14 09:12:50-- http://cakecentral.com/ Resolving cakecentral.com... 174.129.211.41 Connecting to cakecentral.com|174.129.211.41|:80... connected. HTTP request sent, awaiting response... 404 Not Found 2011-01-14 09:13:01 ERROR 404: Not Found.
Once the root page of a site starts returning 404s, we have to start taking guesses about the best way to handle it. Best advice for Cake Central: make sure your root page returns a valid HTML page with a 200 response code.
P.S. If the webhost is trying to do something sneaky, e.g. things work for browsers, but wget or Googlebot is treated differently somehow, the owner of Cake Central can use our free "Fetch as Googlebot" feature in our webmaster console to help diagnose the problem.
Summary: not the weirdest search result I've seen by far. Webhosts that serve up 404s, redirects, or duplicate error pages can cause arbitrary things to happen in search engines. Bing doesn't have the url cakecentral.com indexed at all, for example. Blekko has it, but their page is from Nov. 11, 2010, so they're probably missing the 404 issue by being a couple months older.
alias h='curl -sIw "Time: %{time_total}s\n" -X GET'
This issues a GET on the given URL, printing only the response headers and the time elapsed. Append -L if you want to follow redirects. cakecentral.com. 90 IN A 174.129.211.41
If you look at http://174.129.211.41/ without a host you'll see that it's a nginx reverse proxy/cacheboth beerby and cakecentral are on EC2.
Pretty likely that an error was made at some point in the cakecentral nginx config to include a beerby EC2 private IP as part of a load balance pool or as the single back end (either fat fingered or by retaining an old IP as instances were stopped and started).
It has since been corrected (probably?), but as cakecentral.com is returning a 404 to robots on their homepage the best, most recent return google has was when it was misdirected.
http://grab.by/grabs/32ac4e9cade57bedcc96c8e42fb66a2f.png
DA = Domain Age
PR = Page Rank
IC = Indexed Content (pages)
BLP= Backlinks for the page
BLD = Backlinks for the domain
BLEG = Backlinks from .edu/.gov pages
DMZ = Listed in DMOZ
YAH = Listed in Yahoo Directory
Title, URL, Desc, Head = Whether the keyword is included in any of those
CA = Google Cache age
Screenshot of the screen from which the table has been taken: http://grab.by/grabs/323101a2a3382f4c75b2f077a481931c.png
Disclaimer: I know the guys from SEOmoz fairly well, and have used their site for years.
In an interview with the Washington Post in 2006 Marissa Mayer from Google said that almost no one ever uses the "I'm feeling lucky" button:
http://www.washingtonpost.com/wp-dyn/content/article/2006/10...
But maybe in 2011 Google users are luckier.
because here we see a beerby page with a cakecentral URL http://www.google.com/search?q=site%3Acakecentral.com+beerby...
i would guess it was either a server (housing) accident or a DNS f*ckup that let beerby and cakecentral switch places (in an erroneous state) for a short time, bad thing google picked u the cakecentral home page URL in that moment. it saw it as either a redirect or a direct douplicate of the beerby site and decided to show the older indexed page with the same content (the beerby error page).
yeah, either this or google screwed up.
update: why i guess this is because i have seen similar errors when sombody screws up redirects from the home page. (makes HTTP 302 redirects from the home page to another page, and that page (or the redirect) is then changed to something else...) but this is the first time i ever see such an error between two unrelated sites.
edit I'm wondering if maybe something F'ed up in Google's database
Similar cases have happened before. There is a forum by Google for Webmasters where you can tell Google about problems with your website:
http://www.google.com/support/forum/p/Webmasters?hl=en
You could tell them your findings and maybe someone from the Google team will look into the matter, if you are lucky.
These terms only appear in links pointing to this page: cake central
Looks like the good old Google bombs still work :) http://en.wikipedia.org/wiki/Google_bomb
the note
"These terms only appear in links pointing to this page: cake central"
always shows up as soon as the query words could not get found on the cached page.
On the other hand, we can never be sure if pages exist or where they are on the web that link to our pages with a certain anchor text. The link: operator is broken since a long time and shows only a small subset of the pages linking to the page in question if anything at all.
A more complete list of links can be found in the Google Webmaster Tools, but this is also never 100% complete or up to date. And we can use the Site Explorer to get on the quest to find a certain link:
or lets phrase it like this
there is absence of evidence that it was a link bomb
That said, I have no idea why that page would rank on those terms, error or not.
EDIT: even better - searching for cakecentral.com also leads to the same error page on beerby.com
Edit: It worked at the second attempt beerby.com is hosted at Acquia hosting and cakecentral.com at Amazon
In popular french "cake traces" refers to brown marks in underpants. I guess the french expression "cake face" is a subsequent derivation from it. So I'm trying to guess what "cake central" might mean ...