Understanding web pages better
googlewebmastercentral.blogspot.com
googlewebmastercentral.blogspot.com
2008: http://moz.com/ugc/new-reality-google-follows-links-in-javas...
2009: http://www.labnol.org/internet/search/googlebot-executes-jav...
2011: https://twitter.com/mattcutts/status/131425949597179904
2012: http://www.thegooglecache.com/white-hat-seo/googlebots-javas...
From the 2012 article: "Google is actually interpreting the Javascript it spiders. It is not merely trying to extract strings of text and it does appear to be nuanced enough to know what text is and is not added to the Document Object Model."
Content augmented with JS -- OK to index. Blank page with all content rendered by JS -- not indexed
And what they're saying today is: "we are now going to be able to index single page apps that don't have any server rendered content"
Perhaps I've missed something, but this is my interpretation: SPAs are now full citizens in the SEO world.
Certainly somewhere in the 2008 to 2009 timeframe we saw that the Chinese language version of the Wall Street Journal had a lot of pages in their archive where all of the non-boilerplate content was rendered via JavaScript. Since they didn't do this with their English content, it didn't seem to be an attempt to hide content from search engines, but much more likely a workaround for an older browser that wouldn't properly render Unicode, but with a JavaScript engine that would properly render Unicode. Sometime in the 2008 to 2009 timeframe, Google's indexing system started understanding text that was written into documents from JavaScript body onload handlers, and the Chinese Wall Street Journal archive content was exhibit A in my argument that my changes should be turned on in production.
I'm sure they've increased the accuracy of the analysis since, but they've certainly been able to index content written by JavaScript for something like 5 years now.
Edit: without giving away any Google secrets, here's a pretty good analysis of my work from 2008: http://moz.com/ugc/new-reality-google-follows-links-in-javas...
Edit 2: Since the 2008-2009 timeframe, Google also notices when you use JavaScript to change a page's title. I caused a crash in Google's indexing system when I made a bad assumption about Google's HTML parser's handling of XHTML-style empty title tags <title/> and tried to construct negative-length std::strings from them. When your code runs on every single webpage that Google can find, you're certain to hit corner cases you didn't anticipate. I did test for empty <title></title>, but not <title/>, and made incorrect assumptions about the two pointers I'd get to the beginning and end of the title.
Thanks for sharing, and nice work btw.
V8 and Chrome weren't even a glimmer in Google's eye back in 2006, so I hope they've largely replaced the code I was working on with something based on Chrome. As late as 2010, the DOM was a completely custom implementation that looked somewhat like Firefox, but with enough IE features to fool lots of other pages that would otherwise change their content to "You must run IE to view this page". (On a side note, as much as many people would like to see such pages heavily penalized and indexed as if the IE-only message were their only information, some of those pages are unique sources of invaluable information and users wouldn't be well served by such harsh treatment.)
Google seems to be ignoring the title changes our JS makes. Should we not have the title tag there in the first place and then add it in when the title is known?
This "just a tool" lets webmasters see how google sees their pages, which means SPAs are becoming safe for SEO. That tool plus the song and dance about understanding the modern web is a pretty strong signal in an industry that operates heavily on rumors.
Hopefully this news is about them improving that two week lead time so that more modern websites have a better chance of being indexed.
Source: I worked mostly on JavaScript execution in Google's indexing system from 2006 to 2010. I don't know the intricacies of the crawl scheduling, but I do know it's highly non-trivial.
I'm glad that they included this.
I get that Javascript is required to make certain sites work the way they do, but I'm appalled by the number of sites that require Javascript just to display static text.
Google themselves are guilty of this. Google Groups is (for the most part) just an archive of email mailing lists, but try reading a thread on Google Groups with Javascript disabled![0]
There are very few sites that cannot gracefully downgrade to at least some degree, and there are very good reasons for doing so. A major one is that AJAX-heavy sites tend not to perform well on slow connections[1] (again, assuming essentially static content here). If you want your users to be able to access your site on-the-go, graceful degradation is your friend.
[0] It's especially ironic now that Google Groups is the only place to read many old Usenet archives going back as far as the early 1980s.
[1] Try browsing Twitter on a slow (ie, tethered, or "Amtrak wifi" level connection). For a website that originally originated as a way to send messages over SMS, and is still used that way in other parts of the world, it degrades amazingly poorly over slow connections.
The cynic in me wants to say that this mentality is pushed for by companies like Google because no Javascript means no spying. But honestly I think it really just comes down to laziness. So few people truly care about their craft.
Edit: And since I replied to a small part of your comment, I should say that I disagree completely with your "few people truly care about their craft" statement. At least, I think that writing code that handles a lack of Javascript is only valuable if you have enough users to justify it. i.e. if you spend 20% of your time working on features for 0.1% of users, then you are doing a disservice to the rest of your users. Even more so if you have to compromise the experience for everyone else such that degrading is an option.
In some cases, you go out of your way to accommodate small fractions of your audience. ARIA and catering to those with disabilities is a good example. But turning off JS is a choice; one I respect, but feel no obligation to cater to. I think pages should show a noscript warning, but other than that, its a matter of engineering tradeoffs.
I encounter these (a lot of them blogs -- a perfect example of static text, maybe images, that should be readable with just about any browser) when searching via Google, and the text-only cache option tends to be quite useful for getting the text that I want to read. If that doesn't show the content, then I go back --- there's plenty of other sites out there, if you don't make it easy to read your content I'll just go somewhere else where I can find the same thing.
At least with Google Groups it seems some of it is readable without JS now -- e.g. try this link with JS off: https://groups.google.com/d/forum/comp.lang.python
Google used Chrome to generated page preview pictures (at index-time) to show the search terms (mouse over, but this feature is no more, as it seems). Some websites that shows you the user agent displayed the Chrome user-agent in the preview picture, back then (~2 years ago).
it was called Instant Preview: http://googlesystem.blogspot.co.at/2010/11/google-instant-pr...
details: https://sites.google.com/site/webmasterhelpforum/en/faq-inst...
Google removed this useful feature in 04/2013 :(
As we’ve streamlined the results page, we’ve had to
remove certain features, such as Instant Previews.
-- https://productforums.google.com/forum/#!topic/websearch/Aom...I always thought this was the secret reason Chrome was built. Build a better Googlebot and then, wait a minute, why not just release an awesome browser to get more people using our product at the same time? Forked.
If the Googlebot used a web engine that wasn't found in a common browser, then they would have to pay the costs to make it compatible with all web pages out there. Instead, they push that cost onto developers targeting Chrome. All work to make websites work right in Chrome also make them work for the Googlebot, effectively standardizing the web onto something that the Googlebot knows how to process.
Are Google going to follow their own advice here? Try visiting the official Android blog with Javascript disabled
http://officialandroid.blogspot.co.uk/
In fact, try visiting a whole bunch of *blogspot.co.uk sites with Javascript disabled and see how "gracefully" they degrade. Remember, these are blog sites with mostly text content. And yet Google won't serve them up without Javsacipt enabled.
http://officialandroid.blogspot.com/2014/04/new-mobile-apps-...
And you will see the text content. It sucks, anyway.
If Google can actually index and rank a businesses website that is, say, pure BackboneJS that would be awesome. But I'd like to see it in the wild before trying to sell something like that.
For example, AirBnB appears to be using Backbone here: https://www.airbnb.com/s/San-Francisco--CA--United-States
Is Google able to crawl their listings by executing all the JS? Or is AirBnB implementing other tricks to get indexed?
* Google Now Crawling And Indexing Flash Content (2008): http://searchengineland.com/google-now-crawling-and-indexing...
* Adobe page: https://web.archive.org/web/20080702135702/http://www.adobe....
Adobe is working with Google and Yahoo! to enable one of
the largest fundamental improvements in web search
results by making the Flash file format (SWF) a first-
class citizen in searchable web content.
Google uses the Adobe Flash Player technology to run SWF
content for their search engines to crawl and provide the
logic that chooses how to walk through a SWF.
Edit: parent commenter edited/changed his text quite a bit, originally it was about Flash contentRan came up with an API for the hooks he needed that didn't give away too much of the most clever parts of what he was doing. The belief was that Google would get the hooks it wanted and in return, Adobe could share the special Flash player with other major search engines and everyone would be indexing Flash content. The hope was that Google would just be doing it a bit more cleverly than the competition. I'm not sure if any of the other major search engines ever used the hooks Ran designed.
Source: I worked on Google's rich content indexing team from 2006 to 2010. Ran worked mostly on Flash indexing and I worked mostly on JavaScript indexing.
If they crawl your page with javascript enabled, and find that after a hover event a button appears and after a click on that button a modal appears, and that modal has content about BLUE WIDGETS, they are still NEVER going to rank that URL for "BLUE WIDGETS"
Google wants to send users searching for "BLUE WIDGETS" to a page where content about "BLUE WIDGETS" is instantly visible and apparent.
"If your web server is unable to handle the volume of crawl requests for resources, it may have a negative impact on our capability to render your pages. If you’d like to ensure that your pages can be rendered by Google, make sure your servers are able to handle crawl requests for resources."
I figure Google has to have some form of safeguard against this. Either CPU or network bandwidth limited, likely a time limit too.
In my head I'm picturing a crawler locked in a loop of forever querying random google search results and adding them to a page.
Across individual requests, they have distanced themselves from setting a concrete limit to request sizes[1] (they used to only cache the first 100Kb, then people saw them caching up to 400Kb, now they definitely index things like PDFs that are much larger.)
Obviously the claim probably isn't literally true, but it they could certainly share a JS engine (V8) at the very least, and the idea that the motivation behind developing that JS engine may have been spidering JS-dependent pages doesn't seem too far-fetched.
This instantly puts me back on the Angular hunt, as now I don't have to pay for a service as ridiculous as 'static page SEO'.
The thing is, Google is not a single entity, like any other big corp. So there will be teams doing things differently, even in conflicting ways. Google is not using Angular for most of its sites, I think Closure is more popular there.
You can do static page SEO with JS using libraries that allow isomorphic processing, see rendr, react etc.
Not necessarily. For instance, I would talk about a "database snapshot", meaning a "backup of a database at a particular moment in time".
Obviously the word "snapshot" originates from photography, but to me, the connotation is much more "capture an exact copy at a moment" than "make an image".
The web is fundamentally a document retrieval system. Content needs to work without Javascript.
I just noticed a comment on HN that I wrote a few minutes ago, was already in Google search results - was shocked a bit.
So this isn't really new or surprising