Does Google crawl dynamic content?
centrical.com
centrical.com
As far as we can tell, it makes no difference from if it was generated server side: https://www.google.com/search?q=site%3Aappapp.io
So yes, Google definitely does index dynamic content. I would love to know if it ranks it equivalently.
Also, Bing does not: http://www.bing.com/search?q=site%3aappapp.io
(apologies for the minor self-promotion)
Our goal is not to have every app indexed (as that will by definition be non-original content), but to have our app category pages indexed, e.g. https://appapp.io/gb/genre=Games;has_iap=false;price=Paid/se...
BTW, if you don't want every app to be indexed, I would recommend you the tag <meta name="bots" content="noindex"> in the app pages. Alternatively, you could define the canonical URL as the URL of the original content.
The only thing we cannot seem to get right are the meta title and meta description. If you set that asynchronously based on the React page you are rendering, Google only seems to pick it up in about 10% of the pages. So the SERPS doesn't look as pretty as you would like. I didn't found a solution for that yet. :-(
Now that we know we are (by Google at least) we can put some focus on optimising our SERPS.
I recently did a talk at GDG about the efforts to get our SPA ranking in the Google SERPs, maybe some enlightening parts in there :)
TL;DR: All possible, same rankings, some caveats though.
GDG DevFest 2015 - We can't use Angular. It will hurt our SEO. [video]
Google should give the actual article URL a higher score and pagination pages a lower score. So that in their search results I see the content first and the "dupe content" on pagination pages not at all (or way down). (at least for common blog software)
Now imagine that Google links to example.com/?page=2 as it found the search phrase also there (at a given time only Google knows). So when the user clicks on the search result link that leads to example.com/?page=2 NOW should the blog software know what Google or the user wants?
One thing that comes to my mind is to use the referrer and if it's a common search engine parse the s=SEARCHTERM string and use an internal article search to find the best matching article.
Also, sneaky web sites often give different results to the googlebot user agent than to a non-google firefox user agent
https://en.wikipedia.org/wiki/User_agent
https://addons.mozilla.org/en-GB/firefox/search/?q=user+agen...
Google has "only" 70% market share, so it seems irresponsible to make engineering decisions without testing the others. Google+Bing+Yahoo+Baidu get you to 98%.
Trying to find any of the other search strings in the article for the different loading variants does not return any results. So no variant of javascript injected content is working on Bing currently.
> So, very soon, the days of pre-rendering PhantomJs snapshots and serving shadow content to spiders will be over.
To be clear: webmasters of sites with dynamic content should not celebrate yet. There are still influential spiders other than Google's that do not parse JavaScript (for example, Facebook[1] and Twitter[2]).
[1] https://developers.facebook.com/docs/sharing/webmasters/craw...
[2] Can't find an official statement on this, but https://twittercommunity.com/search?q=javascript%20crawl
The content may be indexed, but if your visitors are on a mobile network, that initial visit (or a visit with stale cache) is going to be crappy. It's great that they can read in they content (though bing cannot), but if it's buried on page two, does it even matter?
As someone who is a proponent of web perf, these kind of articles make me worried that server side rendering will be ignored because "SEO works now for Javascript", even if it's slow and google is only 70% desktop & 80% mobile search.
Though I don't think it's happening, I've thought it'd be very clever if users became the search spider for Google, telling them when content had gone stale and/or doing the spidering on Google's part. Just by using Google's browser.
Now, that's only one bit of data but if you want to be sure you can set up a trigger page of your own.
The way it works is simple, I made a random url, stuck a script in there that sends an email with the url as the subject header. The first script I visited using chrome, the second script I mailed myself a link to from another email account to a gmail account. Both scripts fired when chrome activated the links so I know they work, then I simply let it rest.
I guess the gmail one would require re-crawling of all gmail for it to fire, the chrome only one would be dependent on the version of chrome that I ran the test on phoning home with the link, and I did not re-try this for every version of chrome (or on every platform). So it's not a perfect method but it definitely puts the lie to chrome or gmail data being directly used to power google search results that people would expect to remain confidential because only they have the urls. Score one for obscurity and nice of google to ignore these paths.
I could set up another url for a test using the DNS but you're free to do so yourself as well of course, it is definitely an interesting idea.
And if either of those scripts ever does fire I'll rip google a new one, that would be the sort of abuse of trust that gets my temperature up. But for now I'll take the fact that it hasn't happened as proof that google can be trusted with data to some extent.
This includes navigating the URL structure (in finding blah.com/one/two/three.html there may be attempts to /one/two and /one) and password-protected admins which are not linked to from anywhere (we suspect Analytics or toolbars are telling Google these pages exist). As a result, Googlebot generates a bunch of false positives in our error logs.
Also, maybe it's complementary to the headed version that lot of persons use and reports back. That millions of persons requesting and rendering html, are some kind of distributed indexers too.
Are there other headless browser beside PhantomJS? PhantomJS is based on webkit (Safari).
So a crawler based on a headless browser that consumes little memory and runs for weeks is a major achivement.
It would be also interesting to see what timeouts it still allows. I wouldn't be surprized if the modified browser "virtualizes" time and runs window.setTimeout immediately. Maybe you could make a busy loop and find out what the real timeouts are. It seems there got to be some, otherwise this would open a way to DOS the crawler (not that I'd do that).
http://searchengineland.com/tested-googlebot-crawls-javascri...
But are there any experiments/results related to SEO impact/crawl frequency etc?
Recently Google changed their search result page for tablets. First it looked fine, and useful.
But many times the first result page is now completely full of advertisements, only the second page now shows usual links to websites like Github, Wikipedia, Youtube, etc. of a common search term. Very annoying! And the Youtube link is broken on iPad (it tries to link to a non HTTP address). I am just unlucky to be part of an AB-testing?
An news article about the changes: http://searchengineland.com/google-launches-new-search-resul...