A 2x Faster Web
blog.chromium.org
blog.chromium.org
Interesting parts:
- "...make SSL the underlying transport protocol, for better security and compatibility with existing network infrastructure." Won't this break many caching models?
- "...provides an advanced feature, server-initiated streams. Server-initiated streams can be used to deliver content to the client without the client needing to ask for it." Nice, push without comet-style hacks.
EDIT: more items that stand out
- "SPDY implements request priorities: the client can request as many items as it wants from the server, and assign a priority to each request"
From the protocol document: http://dev.chromium.org/spdy/spdy-protocol
- "Content-length is not a valid header"
- "Clients are assumed to support Accept-Encoding: gzip. Clients that do not specify any body encodings receive gzip-encoded data from the server."
- "The 'host' header is ignored. The host:port portion of the HTTP URL is the definitive host." I guess this supplants the need for SNI support. Edit: No it doesn't since SPDY sits above SSL. The host isn't known until the secure link has already been established.
It is human nature that people don't do things that sound like hard work until they have to, and the continuing use of IPv4 with NAT and other hacks falls firmly into that camp.
(1) Every device becomes addressable again - I remember when it was normal to assume that devices could be reached directly, and would be fire walled if required. That led to a much greater number of people running services from their machine. From the perspective of a startup the idea that a client can run data services without some horrible <nat-related> hack is really interesting.
https://www.londonfgss.com/thread32727.html
It's just a community forum, but things like ACTA and the erosion of privacy that our government is gradually executing all make me really question whether there are downsides.
I'd originally made it available over SSL to help people access it from work without the proxies blocking it (works very well btw), but now it'll be everything.
As part of this I'm doing a review of the styles, images, javascript, etc to ensure that I perform as few requests as possible due to the lack of caching.
The software I'm using hasn't really considered requests. The view seems to have been "it's a one time hit and then it's cached", whereas in the SSL world you'll be good if you get a cached once-per-session model.
It also complexifies things on the back-end, thinking of things like software load balancing and pass-through requests.
All good stuff though, I get to learn about how to make this stuff work and scale at the same time as providing a great feature (privacy and security) to my users. Traditional models are worth breaking to gain these benefits.
The reason HTTP has been so successful has been it's simplicity. HTTP with it's human readable format is not the most efficient, but this makes this it simple. Things like header compression to me are missing the point.
SSL encryption will break all caching, and add latency in terms extra round-trip times.
gzip encoding has minimal value on media data types (images, movies) that represent the majority of the data.
This just sounds like a whole load of complexity for a relatively small one-off hit.
It won't break all caching, just intermediate caches. True this does raise bandwidth usage and initial page load latency (assuming static content is not already cached on the client), but to say it breaks all caching is off.
That said, static image serving is not going to benefit from this new protocol all that much. It seems, like much of Google's other work in this area, to be focused on improving the performance of web-based applications.
They specifically say their goal is to minimize latency to make web based applications more responsive. They aren't interested in making you able to download movies faster, what they want to do is make the web more responsive so that they can develop more apps for the web that traditional would have required a desktop client.
As far as I know, Firefox and maybe other browsers don't store on disk content delivered via https by default, which means that final caches will also be affected. Sure, the default can be changed, but it's an extra security risk. If it wouldn't be, the default setting wouldn't be like this.
Update their PageRank algorithm to take into account website speed, and then publish this as a fact.
The idea would be that they would have GoogleBot measure the latency, and overall page load time and rank sites higher that respond faster, and slow sites lower.
One of the problems with PageRank is that most of how they measure site quality is a mystery, but making it plain that speed is used in ranking would bring website performance to the front of everyone's minds. They could use similar metrics as YSlow and PageSpeed use, and tell everyone to optimize their sites using those tools as a starting point.
It would be like comparing the drive to a particular location using completely different starting points.
Plus, people would start doing all sorts of tricks to boost their ranking. Serving static, cached files for Google and giving everyone else a slow, dynamic site. It's way too easy to abuse.
Google just needs to figure out a reasonable baseline latency to subtract - for that, there's ping, ping the previous host on the route, latency to other hosts with close IPs, GeoIP+a database of baseline latency to sites in the same area etc.
And the tricks will be dealt with the same way as now: If you're caught doing them, you're blacklisted. But this will be less of an issue in this case: If you've taken the trouble to set up caches that you can direct Google to -- why not direct your actual users to that too?
The cache trick is that Google doesn't need to see dynamic content as much as the user does. Example: Any site that's based on user selection to determine customization. I might need to see recommended content that's based on a very complex algorithm that takes a while to generate the page. Whereas Google just needs to see a page. Cache for Google, generate from scratch for everyone else.
The cache trick is the same for not-logged-in users as well as for Google. If you're showing the same page to more than one user (and you should be, for your not-logged-in users), you should cache it and show everyone that. If show Google a few hours old version (that you wouldn't want users to see) you'll be punished with the current rules.
If you're only showing content to logged-in users, Google can't see anything, anyway.
So, in theory, the faster your site responds, the more pages googlebot can crawl without crushing your servers, the more people find your site when searching, thus giving more people the opportunity to visit your site because it's fast.
Regardless, Google is doing this to push at the technical limitations on apps like GMail or Google Reader, not Joe Shmo's burger shop.
IMHO a site that cares about the user experience would have been optimized to load quickly.
Heck, even if Google told people speed was important, and didn't incorporate any speed measurements into PageRank, I think it would have a positive effect.
I think that site speed does correlate with a better user experience. A site that loads quickly shows that the person who made it cares about my experience more than a bloated site that takes forever to load.
I will admit that it's certainly not as as strong an indicator as the content, inbound links and other factors that are rumored to be part of PageRank but speed does matter. All other things being equal, I would much rather visit a site that loads in 2 seconds than one that loads in 20 seconds.
They're the only company with the means to take on such massive, unprofitable projects, and at the same time enough street cred that everyone won't immediately assume they're trying to make the web proprietary. The fact that they tend to open up these things from the get go helps, too.
I think the important thing to note is that Google is the only large company whose business model depends almost completely on the success of the Internet.
This is why we see them at bat for Net Neutrality as well as efficiencies within communication protocols.
Microsoft, Apple, Oracle, HP, IBM . . . those companies make their livings in other ways.
The only company comparable to Google would be Yahoo in this way. Their inability to keep engineering talent means their contribution to The-Internet-As-Platform is not as significant.
IE. Imagine a personal computer market with no way for computers to communicate with each other, would users be still buying computers as much as they do now or would it be just a small niche market?
Even though there are several big players whose business model depends quite a bit on the internet, yet no other tech companies does as much as to improve the internet as Google does. The icing on the cake is that Google opens up and shares most of these projects to everyone for free.
Google may be the biggest company whose success depends upon the Internet, but they're far from the only one.
They are hardly the only company doing this. Microsoft Research do a ton of pretty visionary / far-reaching work, much of which has nothing to do with MS-proprietary platforms. For example, the VL2 architecture for data center networks: http://sns.cs.princeton.edu/2009/10/new-datacenter-networks/
This makes me uncomfortable.
You just don't trust your money to an apparently unarmed bank-robber.
As for licensing, Microsoft owns the said IP and can re-license it on a whim.
Not all of Microsoft is evil, but they, as a company, cannot be trusted to behave in a civilized manner.
AFAIK, the Simons publish research papers freely, and GHC is licensed under what amounts to a BSD3 license (by the University of Glasgow... not Microsoft).
I wonder if Microsoft will one day redeem themselves to the public like IBM did?
I don't regard the shameful approval of OOXML as an ISO standard, and all the questionable things Microsoft did in order to accomplish it, as a particularly nice thing to do. There are a lot of NBs that were formed specifically for that purpose that were never again heard from. As it is now, the damage to ISO seems irreversible.
Also, there is the trolling about "Linux infringes 4 bazillion of our patents" empty threats (because if they weren't empty they would have acted) and the failed attempt to sell some of those patents to trolls that would go after Linux users. The fact if failed (because a non-invited group won the auction and donated the patents) makes it not less evil.
No. Being evil is in their soul.
They may be less evil now - I can agree on that. But it's so much more because they are far less relevant now than they were in the 90s than any particular change of heart.
They deserve no redemption.
They'd now have to pony up for an SSL cert, which at least doubles the website ownership cost, as well as potentially go through a very complex process to get the certificate issued.
It also completely eliminates the possibility of them having a free website on their own domain, unless an easy and free method to obtain an SSL cert becomes available.
The gain to them in these circumstances is negligible, if they aren't doing any e-commerce.
You have the public key via DNS, you send an encrypted request and receive a verifiable encrypted response that includes content and begins the multiplexed channel.
For this to take hold in the wild, it has to be possible for me to run a combined HTTP+SPDY server, and for a client connecting to it to automatically make use of SPDY if it is capable. This must happen without user intervention (i.e. we don't expect people to type spdy:// instead of http://).
It seems like this should be possible. Perhaps there could be a special server response that instantly "upgrades" a TCP session from HTTP into SPDY. But I'm not seeing anything about that in the docs; is this part of the SPDY plan (yet)?
You can already do this with Flash today, and it would be much easier for Google to push it into the DOM than to touch the transport protocol.
But all this is still experimental.
In general I'm more happy with web standards. Google doing research on this is great, but if they unilaterally push solutions that will anyway hit the mass market because they are Google is not good.
All the web, the idea of HTTP APIs and so on are based on the fact that HTTP may not be perfect but it's trivial to use, implement, and so forth. Please don't break this fact.
But how do you gain adoption of such a tech across clients and servers (which was the problem with pipelining) will be an interesting question to see
55% = Reality so far
From http://dev.chromium.org/spdy/spdy-whitepaper
Q: Is SPDY a replacement for HTTP?
A: No. SPDY replaces some parts of HTTP, but mostly augments it. At the highest level of the application layer, the request-response protocol remains the same. SPDY still uses HTTP methods, headers, and other semantics. But SPDY overrides other parts of the protocol, such as connection management and data transfer formats.