JQuery 1.6.2 syntax error? You may be the victim of SEO.
encosia.com
encosia.com
I work at Google as a Webmaster Trends Analyst to help webmasters with issues like this one.
Looking into this, the first thing I noticed is that the blog.jquery.com seems to be blocking Googlebot from fetching its pages, but the site responds normally for web browsers: it returns an HTTP 500 error headers for requests using a Googlebot user agent. You can see this yourself using a public tool like Web Sniffer to fetch the page spoofed as Googlebot ( http://web-sniffer.net/?url=http://blog.jquery.com/2011/06/3... ) or using Firefox with the User Agent Switcher and Live HTTP Headers addons.
Unfortunately this is a very common problem we see. Most of the time it's a mis-configured firewall that blocks Googlebot, and sometimes it's a server-side code issue, perhaps the content management system.
Separately from that, I also notice that the blog.jquery.it URL is redirecting to the blog.jquery.com, suggesting they are fixing it on their end too.
If an jquery.com admins want more help, please post on our forums ( http://www.google.com/support/forum/p/Webmasters?hl=en ).
Cheers,
Pierre
Thanks for the details on this! We dug into our Wordpress and realized that W3 Total Cache was configured to block anything that had the word 'bot' as part of its user agent string (sigh). That's now fixed and live.
As to the redirect - that's actually a bit of devious magic on our end. Since jquery.it is hotlinking to our JavaScript and CSS we just added a bit of JavaScript to automatically redirect them to the right jQuery.com page:
if ( /jquery\.it/.test( window.location.hostname ) ) {
window.location = window.location.href.replace('y.it', 'y.com');
}
Thanks again for your pointers - looks like the jQuery 1.6.2 blog post is already showing up in the search results and jQuery.it is not in the first page of results (although "jquery 1.6.1" hasn't updated yet and jQuery.it is still there - I assume that it'll be remedied as the Google spider makes more progress). Thanks again!Yep, that'll do it :) Glad you found it and fixed it quickly!
One follow-up thought for you and anyone in this situation: Set up a Google Webmaster Tools account ( http://www.google.com/webmasters/tools/ ) and make reviewing it part of the webmasters' daily routine. In this particular case, the Crawl Errors page in the Diagnostics section would have flagged this problem very quickly.
A couple of pro tips for Webmaster Tools:
1. Be sure to look at the "date detected" column because it is accurate and somehow people miss it.
2. Set up email forwarding for messages Webmaster Tools sends to you: http://www.google.com/support/webmasters/bin/answer.py?answe...
This is the real problem here, because people do have a lot of trust that very highly-ranked results on Google won't hurt them.
That said, I find it funny that there is so much vitriol against Google in this case when, as has been noted above, it's likely a misconfiguration on jquery.com's part.
I don't think your explanation holds. Pagerank is based on reputation (created by links and other means which Google isn't very specific about), more than on contents. jquery.com has a pagerank of 8, so it can't be all inaccessible to Google. The GP says that blog.jquery.com is inaccessible. So how does jquery.it get a pagerank of 8? Your explanation would look correct if jquery.com had a pagerank of 1 (due to misconfiguration) and jquery.it sneaked in with a rank of 2. But this isn't the pattern.
However, this still doesn't explain why jquery.it has such a high pagerank. And blog.jquery.it has a pagerank of only 1. I can't find any interesting-looking links to jquery.it that would explain the high rank. It's also strange that Google returns blog.jquery.it 1st for the search "jquery 1.6.2", since those strings are found at lots of highly ranked blogs (pagerank>1 for sure). We could be dealing, once again, with Google's preference for finding search terms in the domain name. But this puzzle is still not coming together for me. This is why I'd like wiser persons than me to try to get to the bottom of this.
EDIT: 1 more datapoint: bing doesn't return anything from *.jquery.it for the search in question.
Just checked, I'd gotten it from jquery.it...
Eg) I have one for /objc that searches a half-dozen high-quality, good signal-to-noise, non-spammy objective-c related sites, so that if I search for "timers /objc" I get only quality content and no farmed spam.
There's a lot more to it but they've essentially farmed out the job of whitelisting the non-spammy parts of the internet to their users.
Edit: and more to the point of what you were desribing, you can easily use other people's public slashtags, and it will detect and suggest relevant ones as you use it. It's totally worth playing around with.
Guess what you get when you go to http://duckduckgo.com/?q=jquery THATS RIGHT, an official site logo, because Gabriel is f-ing awesome and I love DDG's little almost insignificant features like showing you the official jquery website vs what you THINK the official one is, since google never helps you there.
Also I have adblocking and opt-out from google's ad tracking on so I never suffer these things. But that's what makes DDG so amazing, that opt-outs don't mater Its just so clutter free. Putting things into context vs just presenting you with data. Note that in my search results I even get the nice icons indicating if the result is spam or not. That website for the fake jquery is... well its not even on first 100 search results, may be blocked.
Thanks DDG, you just justified your existence yet again.
Edit: Yeah, I think that's the case. These searches don't show the "Official Site" badge: http://ddg.gg/?q=jenkins+ci http://ddg.gg/?q=hudson+ci, but these do: http://ddg.gg/?q=jenkins+software http://ddg.gg/?q=hudson+software
Compare those queries to the name of the wikipedia pages: http://en.wikipedia.org/wiki/Jenkins_(software) http://en.wikipedia.org/wiki/Hudson_(software)
Edit: the headline should be: "Everybody watch out, a fraudulent jquery website ranks higher in google than the official website". The syntax error is the best thing that could happen.
We developers should heed his warning to be careful about download sites. AND Google should do a better job of blacklisting spam sites like this.
Doesn't seem that hard to me, compared to the other tasks they do.
Dude fell for what amounts to a phishing scam. Sure he should have been on better guard, but the circumstances definitely contributed to his user error.
"Dude was careless" is the entire point. Google should not have ranked that site above jQuery's site, but you should check that you are at least downloading your JS from the right domain.
Google is providing a service that vouches for the authenticity of sites by their ranking in the search results. They failed.
They try to be helpful by ranking them in some fashion, but at the end of the day it's up to you how you use the results that are returned.
We can either rail against the realities of human nature and persist in blaming the user, or we can accept that asking the user to check everything always is a plan guaranteed to fail, and build better tools to help eliminate the problem.
I vote for pragmatism and better tools.
If the "wrong" one had the same content as the right one, would you know which one to pick?
Only not: http://en.wikipedia.org/wiki/Tenerife_disaster#Safety_respon...
It also has cloned the subdomains: http://dev.jquery.it/ and http://forum.jquery.it/
What this site appears to do is mirror the content of jQuery.com by copying everything and then appending the "Time to generate" string. I just checked adblock, it also adds a Google Ad, which is the point to this.
Obviously Google has messed up big time, but also the whole web by linking so much a fake site that it has the same page rank as the original.
1: http://wayback.archive.org/web/20071001000000*/http://www.jq...
2: http://wayback.archive.org/web/20090601000000*/http://www.jq...
That's probably what happened here. The jquery.it domain has nowhere near enough link strength to get an 8/10:
http://www.opensiteexplorer.org/jquery.it/a!links
jquery.com, for comparison:
and http://code.jquery.it. It's an exact mirror of jquery.com, the only differences are the tld, the borked download file and ads.
They should ban not just jquery.it from both natural rankings and AdSense, but every other site on the same AdSense account and with the same registered domain owner.
Although altering jQuery to add a link at the bottom of every textarea field...
IMHO Google is very vulnerable to competition by new brand-name lookup services.
if(document.location.host != "jquery.com") {
document.location = "http://jquery.com";
}If the canonical tags were added stealthily he probably wouldn't notice at all until his SEO was completely destroyed.
There needs to be an international douche law to serve as a deterrent to this kind of behaviour.
Why this was done? Here were my first few thoughts:
1. Display ad revenue. - Maybe initially, there is Adsense markup but the ads aren't showing for me so perhaps Google has disabled them.
2. Affiliate income from the links to jQuery books - I can't see an affiliate code in the links so probably not.
3. Hijacking the Donate button - No. This leads to a blank page with just a Time to Generate snippet.
Pardon? If you download just any jQuery without even checking the domain you are downloading from, then you are very careless. That's just like typing your Paypal password into a form on a website that was linked in an email that looked like it came from Paypal...
Your copy of jQuery will be able to see anything that happens on the site you are writing, send any user password to a external server, read session keys, query your API for any data as a logged in user, etc. You could even build a botnet out of modified jQuery libraries.
Whenever you download executables, make sure you know where they are coming from!
* if a page A refers to the same external css and images as page B on another site, and those external resources are local to B, then assume B is more original and should be ranked higher than A.
Of course the SEO people will get around this by making sure they take copies of the css and image assets as well as the html ones once this is implemented, but at least it'll save the "target" site a little bandwidth.
Which really is not the issue at all. Do you also suggest that people fearing home invasions paint their walls red, so the blood splatters are hidden if they get shot?
Unfortunately the real problem Google is unable to do much about, aside from a few high-profile things (jQuery would count as high enough profile, but many similar libraries would not). How do they know, given two apparently identical chunks of content, which is the original source?
This is a very naive view of Googles algorithms. While link count plays a significant role in ranking, there are many more factors (including various secret sauces) that determine where a given link will end up in the SERPs.
Still, it really seems odd that Google shows a fake as the top result here. I just ran the search myself and while jquery.COM ranks first for a generic 'jquery' search, the release blog post doesn't appear at all when searching for 'jquery 1.6.2 released'. The closest you get is a link just to 'blog.jquery.com' with the title "jQuery:". It makes me wonder if something went wrong with the blog release that affected the way the "Google Juice" flowed down to that post, and this fake site managed to capitalize on it.
The .IT in question site is the fourth.
Things move fast on the interweb, I can't agree with anyone claiming this is Googles fault.
With the size of the database, the breadth of queries that are done against it and the myriad of possible returns - how could they reasonably police it?
There is an adress field you know.
Frankly, had you been more careful, worse could have happened down the road. Still, I would be interested in seeing exactly how jquery.it made it to the top of the search listings.
Diff suggests the plan was probably "Serve adsense ads." That is the punchline to quite a bit of spam.
Beware of downloading JQuery from jquery.it, which may appear before jquery.com in related Google searches. The end.
No, the victim of a scam and Google. The site barely has any inbound links, if this is SEO, they suck.
This is a google fail, not an SEO fail.
It's like eating a kebab dropped in the street.
jquery.com does NOT appear to have a fully valid SSL certificate: Chrome gives me "the site's security certificate is not trusted!"
Like it or not, Google is an important part of establishing reputation -- that's what pagerank was built on initially and if that becomes worthless then finding the true source of something becomes very difficult.
Hypothetically supposing that jquery.com had a lovely little green lock, that wouldn't matter, because on jquery.it a) you wouldn't be looking for the lovely green lock and b) if you did look for it, look here, a lovely little green lock and c) you didn't click the lovely green lock to see who it was issued to but if you did d) it was issued to jquery.it, which matches the address in your bar.
SSL solves one problem, really really nicely: it makes it impossible to eavesdrop between the user and a trusted endpoint. It does basically nothing to make sure that the trusted endpoint is the one the user thinks they are interacting with.
When I visited by bank's web site and drill into the certificate details I can at least establish that someone my browser vendors trusts (or someone they trust ...) issued the certificate to an _organization_ called 'Bank of Nova Scotia' in Toronto, not just the domain name.
If I was able to register micr0soft.com then hopefully I would have a hard time getting an SSL certificate issued for it. I know there have been a number of discussions on certificate infrastructure here that show how complex this can become.
Quite well designed really :-)
Other good signs are them being linked to from cdnjs, cached-commons or microjs. If I'm looking to solve a javascript itch I'll first browse these sites to see if there is a popular tool.
Also, if you're looking for jQuery then it's because you've read about it online somewhere. Simply go back and follow the links.
Sounds like perhaps someone was testing something long back (cert was signed in 2009) and just never turned it off.
So how I do it is I look for the community. github is a good place to look. HN is itself a good source of vetting. Google, certainly, but not the first link I find. In fact, when I first heard about jQuery, I didn't assume that the "real" site could be trusted either: if I'm going to install this on my site, and serve it to people who trust me, then it had better be trustworthy.
Now imagine I run a tutorial website, and people come to my site because they trust me, and then they install software they copied from me (or my links), and distribute that to their users. Wow. Kudos to the author: I think it was bad form to blame google here, but the fact that he admitted it all does a lot to reestablish trust.
The situation here was that someone was using one of my samples from the jQuery 1.2 era and wanted to see if it would work with 1.6.2. He downloaded the ".it" copy of jQuery to test it with, got the syntax error when he used it in my sample, thought it was because my code didn't work right with 1.6.2, and got in touch with me about it. That's about where the post picks up at, when I rushed to grab a copy of 1.6.2 via Google and made the mistake of downloading the ".it" copy without noticing.
Now is this an argument for keeping the url bar? It's obviously error-prone, but the other methods of establishing identity don't seem to be there yet either.
Except that this wouldn't be news. And because it wouldn't be news, the blame falls entirely on the author.
The author got fished. Kudos for letting everyone know about it. -kudos for blaming google.
Alternatively, Google could just work on spotting phishing and spam sites.