Reactive prefetch on Google Search: 100-150ms speedup
plus.google.com
plus.google.com
Chrome: https://chrome.google.com/webstore/detail/undirect/dohbiijnj... (source: https://code.google.com/p/undirect/ )
Firefox: https://addons.mozilla.org/en-US/firefox/addon/google-no-tra... (source: http://matagus.github.io/remove-google-redirects-addon/ )
They sometimes run into issues as Google tweaks the search results page.
According to the same author, that was deployed over a year ago but only to browsers which support it:
https://plus.google.com/+IlyaGrigorik/posts/fPJNzUf76Nx
Unfortunately, this was implemented in Firefox years ago but disabled due to a fear-mongering campaign by some self-styled privacy advocates who were quite vocal in sharing their misunderstanding of web privacy:
https://web.archive.org/web/20060126211610/http://weblogs.mo...
EDIT: I forgot to mention the new Beacon API, which is getting more traction because it's more powerful and is fully supported as of Firefox 31:
https://developer.mozilla.org/en-US/docs/Web/API/navigator.s...
What 'misunderstanding'? I don't want people knowing what third party links I'm clicking on. There's no misunderstanding. I understand it perfectly. I just don't want it. I disable 3rd party HTTP referers as well (using the RefControl extension). Sure it's not the only way sites can implement this behaviour, and I'm glad an official way exists... but only so Google will use it and then I can disable it. The argument against having an off switch is basically 'well, they're going to fuck you anyway, so bend over and here's some lube'.
You say 'some self-styled privacy advocates' are fear-mongering, well its because webheads keep implementing insanely harmful features and aren't actively making the web better for the privacy concious. They (we, I guess) are grossly under-served.
The response to link tracking should be "Hmm, how can we have websites ask for this permission, and shut down all these nasty means of doing it?" not "oh boy, these people are tracking people in an ugly way, how can we make this fast?". But guess which one is actually a hard problem.
Do you suppose that most everyday users know (not suspect, but know) that Google are watching every link they click on? A technologically illiterate user base cannot consent.
As for me, I think once a site has seen my exact search term, knowing which of the results I'm clicking on is a small leak and quite useful so that the popular results can be put at the top.
Not everything is black or white.
Given two solutions with equal potential for abuse, why not pick the technically superior solution?
Straw man. You're presupposing the existence of 'ping'. The argument is why, given an observation that web features X and Y are being used to implement contentious function Z, would you want to implement a brand new, even more insidious, web feature, designed solely for doing Z, in the first place? Technical superiority of the new implementation of Z is not in dispute.
> they already know which you click on.
I'm pretty confident that they don't in my case.
Would you mind sharing the details of how you achieve this?
I think in this case it's pretty black or white, because people would ordinarily take the time to add to their search queries. If people want to see results on (whatever fetish), do you think they would just Google the word 'sex' and then go from there, under cover that they might have just been looking up the Latin word for 6 they saw in some inscription? (Whoops, also better load the first 9,700,000 results pages in javascript, which is what I estimate you have to read through to get to something that mentions this without adding the word 'latin'). Wouldn't want to give away which page of results the user stopped on.
I mean if they want 'suicide help in Detroit' would they just Google the word suicide?
Not everything is black or white, but for Google to know what the BEST result is for 'suicide help in Detroit' is pretty black and white: they already have the query, and yes, they should know which link is clicked on the most. If it was originally on the second page through their algorithm but gets 90% of the clicks when they put it on the first page, yes, they should absolutely know and use this information. In fact (because it had been on the second page in this example) it might save lives!
there really is no trade-off or drawback. it's like the server in a restaurant asking you if you enjoyed your meal, and using this as part of recommendations the next someone someone asks what's popular.
"Well, you know what I ordered, but damned if I'll let you know what part of it I liked. That's just too personal."
reads to me that way anyway.
...and then being able to get you fired, steal your identity, or otherwise affect your life without your knowledge because you happen to enjoy a certain kind of food your boss doesn't like.
The metaphor is spot-on. "Did you enjoy your meal"? "Oh so you can get me fired, steal my identity, or otherwise affect my life without my knowledge because I happen to enjoy a certain kind of food my boss doesn't like? No comment."
The reason it's a good analogy is because you already 'ordered' (the search terms are the meaningful data) and a statistical sampling of which of the top 10 results for "octopus hentai videos" a random sampling of users who entered that search term clicked on, is not in practice used for anything other than improving ranking quality.
it's exactly the same as "did you enjoy your meal" or what you liked about it - you've already given up most of the info by ordering in the first place.
so your analogy is a good one. it's just completely innocent.
* Wikipedia (smartphone, PC): Aard Dict[1]
* Translation apps (smartphone)
* OpenStreetMap (smartphone): OsmAnd[2]
[1]: http://aarddict.org [2]: http://osmand.net
Only because people have "forgotten" how to build offline apps.
* it's nice to have knowledge at hand when offline (commuting via train, abroad w/o a local data plan, no gsm coverage, edge connection instead of 4G)
* saves bandwidth for my 1GB/month data plan when appropriate
* Aard dict can be used to evade filtering (not my primary concern, but there are people in other countries who might benefit)
Search is one of these. In order for search to work, the user has to send a query to a third party. The third party has to be able to read the query to do the search. If that's not an acceptable loss of privacy, then don't use search.
This is exactly what I was talking about: the misunderstanding is thinking that your outrage changes the privacy situation in any way. The options on offer are “Stop using Google” or “Let Google collect data about the search results you click on”; redirect scripts, <a ping> and Beacon are all simply implementation details for the latter option.
If you feel strongly about this contact your politicians and lobby for privacy laws restricting the data companies are allowed to collect. There's approximately zero chance that everyone will voluntarily stop measuring how well their search ranking algorithms work.
While you may be right that the privacy angle is something that few users care about, there are real tangible benefits to disabling some of this stuff.
b) It's true that you paid for the device but Google pays for the service – the deal is that they pay for everything by showing you ads. If you disagree with that you're welcome to use or start a competitor but it's absurdly entitled to think that you have any grounds to demand that they change their business model.
b) Don't put words in other peoples mouths. How incredibly rude of you. I have not demanded that they change anything. My post makes no mention of any "business model".
My computer as of today allows me some control over the code that runs on the CPU and the data that gets transferred over the network. I'd like to exercise that option if and when I wish to. You're free to do whatever you want.
As others have said, there is no way to prevent Google from gathering this information, so adding the capability to a browser doesn't change the situation.
Actually, I could never understand, why Google substitutes the link with the redirect on the RIGHT click. It is necessary when you LEFT click, the page looses control and it doesn't matter much for users anyway. But WHY do it on the right click? When the user RIGHT clicks the link, the user needs the ORIGINAL link! And the user IS NOT leaving the page. Google is free to record a right click. And yet, instead of just recording the right click Google mangles the link instead. WHY? Just evil/stupid? or I'm not understanding some hidden wisdom here?
(And BTW, this is a compliment to Google, as I'm asking the question. Not assuming that by default, as any other big co they've just filled with bozos.)
I don't know why, but especially when something is following the default behavior I do not assume malice.
(And BTW, this is again compliment to Google, as I'm surprised. Not assuming that by default, as any other big co they've just filled with bozos.)
// version 2.2.1
// Release Date: 2014-02-28
// http://userscripts.org/scripts/upload/47300
of course userscripts.org is gone now(why?), but you can find this script on other places.
I'm not taking sides here, I don't know enough to make a judgement. But it's interesting that Google seems to be increasing their pattern of standards-tweaking in order to make a superior product - and who can fault them for that? But isn't that how we got so much of the mess that MS made?
What's to be done?
Sites that want to get the same speedup can make sure they don't have any secondary resources that block page rendering.
Finally, the particular mechanism, link rel="prefetch" is used by Bing and IE 11 [1]. Google has just found a way to prefetch even earlier than normal, by inserting the prefetch links into the search page as soon as the user clicks.
[1]: http://blogs.msdn.com/b/ie/archive/2013/12/04/getting-to-the...
The benefit of the web was that you didn't have a fat GUI locally. With the push to aggregate assets together and get the entire UI/logic into the client browser, cached, and read off of remote services, we're slowly making our way back to local GUI apps. This time, the browser is the OS (forgive the poor analogy).
What was old is new again.
Abuse of standards, to me, looks more like intentionally breaking compatibility with other browsers or implementing features that are easy for one party to implement and hard for anyone else to. For example, ActiveX was problematic in part because it was straightforwards to implement on Windows and a nightmare to try to implement anywhere else. This forced other browsers to either make their non-windows users second class citizens, or fall behind on feature parity.
I work at google, but as a ground level engineer working on internal tooling, not on Chrome or Search or anything. My views here are my own.
The early feedback from FF, IE, (and to some extent, Webkit), folks have been positive, and I'm hoping this can be a cross-browser feature in 2015 (yes, I'm an optimist).
I mean it's not great for net neutrality standpoint, but it's a next logical step for them.
For external links, if you aren't a search engine, you don't know what resources need to be prefetched and even if you did you will have no way of knowing when it changes (other then manually). For internal links, in most cases the resources are already cached and http 2.0 server push is going to fix the rest.
Beyond that, the API looks clunky - instead of having a "prefetch" attribute on a link, you need to add more links with special syntax to the header dynamically.
What's the point in adding this functionality to a browser? this looks like an implementation specifically designed to make only google search faster.
you could find out if google set up an API that exposed this information. Then a simple plugin could update the links on your page.
Then there is no latency benefit from requesting in parallel, because there is no round trip to avoid.
The benefit of above technique is that it's deployable today, doesn't require the destination server to be upgraded, and works for cross-origin resources.
HTTP/2, I'm pretty sure, doesn't allow this (and there are all kinds of reasons it wouldn't be a good idea), but it wouldn't necessarily require forcing the other server to push the content.
http://perspectives.mvdirona.com/2009/10/31/TheCostOfLatency...
From Marissa Mayer while at Google:
> Marissa ran an experiment where Google increased the number of search results to thirty. Traffic and revenue from Google searchers in the experimental group dropped by 20%.
> Ouch. Why? Why, when users had asked for this, did they seem to hate it?
> After a bit of looking, Marissa explained that they found an uncontrolled variable. The page with 10 results took .4 seconds to generate. The page with 30 results took .9 seconds.
> Half a second delay caused a 20% drop in traffic. Half a second delay killed user satisfaction.
The article does at least have plenty of examples. I just can't help but think some of these are heavily confounded with "changes cause people to change." That is, I'd be curious to know if any of the tests did nothing other than increase latency.
The increase from 10 to 30 results, for example, had to have changed the look of the page. Would that alone have been enough to change behavior? My hypothesis is that it would be. Especially if there was a viable alternative that was still familiar to the user. Can't fathom by how much, though.
http://chimera.labs.oreilly.com/books/1230000000545/index.ht...
I think Google needs to check its steps quite carefully when doing things like these. For quite some time they have leveraged their search monopoly (think about their EU search market share) to bring search/browser-type integration features to chrome first. I would say this is abusing a monopoly in one market segment (search in the EU) to attempt to create a monopoly in another segment (browsers in the EU) by continually making sure that Chrome is the browser that works better than other browsers when using Google search services.
Yes, this is innovative, but there is also a concept known as antitrust laws. Another way of bringing this to the market would have been to invite competing browsers to use this and build a credible time plan for a simultaneous launch for all the browsers that wanted to support this.
Perhaps their implementation differs, but they even showed the JS they're using to perform the prefetch in the post. It doesn't seem like they're trying to hide anything.
For example, Microsoft performed a technique known as tying.
But even the claim was not simply "they shipped IE with Windows 98" , or that they introduced other product features to work well with IE first, but "they made IE deliberately difficult to remove, and deliberately and intentfully made it harder for netscape navigator to work". This is not the same as "we made IE the best, and did nothing to competitive products". In particular, it is not the fact that they introduced stuff to IE first, it is the fact that they deliberately harmed the other products.
See paragraphs 94, 95, and 96 on http://law.justia.com/cases/federal/appellate-courts/F3/253/...
Even then, the court held the tying should be analyzed deferentially, and that the US would have to prove this had actual anticompetitive effect.
See in particular, the court's admonition at 93
"As a general rule, courts are properly very skeptical about claims that competition has been harmed by a dominant firm's product design changes. See, e.g., Foremost Pro Color, Inc. v. Eastman Kodak Co., 703 F.2d 534, 544-45 (9th Cir. 1983). In a competitive market, firms routinely innovate in the hope of appealing to consumers, sometimes in the process making their products incompatible with those of rivals; the imposition of liability when a monopolist does the same thing will inevitably deter a certain amount of innovation. This is all the more true in a market, such as this one, in which the product itself is rapidly changing. See Findings of Fact p 59. Judicial deference to product innovation, however, does not mean that a monopolist's product design decisions are per se lawful. See Foremost Pro Color, 703 F.2d at 545; see also Cal. Computer Prods., 613 F.2d at 739, 744; In re IBM Peripheral EDP Devices Antitrust Litig., 481 F. Supp. 965, 1007-08 (N.D. Cal. 1979)."
Note that these are also tying between sold products and given away products, not just given away products. Otherwise, open source linux distributions with large market share would have tying issues (and in fact, they've been unsuccessfully sued for illegal competition before)
I'm curious what exact antitrust concept you think this violates.