So that means everyone else gets a slow image format? Yes, but you have saved bandwidth and your pages have loaded quickly for ~50% of your users.
I don't think webp is really a source image format, more something that is browser and Pagespeed specific. Try it, you will be impressed.
Consider this: I have page.html, with an <img src=logo.jpg> tag. If I'm using Chrome and I request logo.jpg PageSpeed/the web server detects that I support WebP via the Accept header. It sends WebP image for the URL logo.jpg.
And that's a problem. Because the content for logo.jpg is no longer dependent entirely on the URL. It the URL, and the values of the User-Agent or the Accept headers, which determine the response.
This is called content negotiation and it utterly nullifies shared HTTP caching. This is bad because shared HTTP caching can be a huge performance win.
The right way to serve WebP images is to change the URL. The application logic that generates the HTML should detect that browser support webp and rewrite the image tag accordingly. In other words, page.html should include a <IMG src=logo.webp> tag.
This means that browsers that support webp get it, browsers that only support JPEG get it, and shared HTTP caches can be used without polluting the cache.
Sadly, PageSpeed and other technologies like it aren't currently smart enough to re-write the URL in the base HTML page. As always, technology that "automatically fixes problem X" rarely is the best solution.
Also, if you respond with different HTML, you still get the caching issue, because you may mistakenly serve cached HTML linking to WebP images to non-capable client. So the cache should still consider Vary and Pragma headers (or it's broken).
The only thing I don't really fancy about "/cute/kitten.jpg" serving a WebP image is ".jpg" part. If content-type may vary, the "extension" part of the "filename" should be generic ("/cute/kitten.image") or missing ("/cute/kitten"). That's for saving pics, when URLs transform into filenames.
First, "Vary:Accept" and Vary:User-Agent effectively nullifies shared caches. Because instead of saving 1 copy of a response for a URL, and serving it to all requests, a shared cache has to save 1 copy per unique combination of URL and the Accept request header and/or the User-Agent header sent by the browser.
Browsers send wildly different Accept header values, not only across browser vendors but across versions as well. And the User-Agent string has insane variation since OS, CPU, language, and more can be included.
The net result is that all these unique request header values fragment the shared cache so much that you don't get any cache hits. The shared cache has been nullified.
Second, if the HTML is a static file, yes you have the same problem. However, in the modern age of web publishing, the HTML is almost always being dynamically generated by some application logic/CMS/something. This means that it is rarely is ever cached or marked as cachable to shared caches.
More about Vary and shared caching: http://zoompf.com/blog/2010/03/the-big-performance-improveme...
http://blogs.msdn.com/b/ieinternals/archive/2009/06/17/vary-...
Unfortunately, “done properly” excludes things like nginx's cache module, a number of CDNs, and even certain browser caches. Some of the bugs are obvious – e.g. returning gzipped data to a client which didn't request it – but others will simply require monitoring to realize that your cache isn't caching anything or has a dismal cache hit rate because every permutation of request headers and the values referenced in Vary are being treated independently. Worse, all of these can change unexpectedly due to updates in “stable” software so you need deep monitoring checks to ensure that everything is still working the way it was when you set it up.
The more I've used content negotiation, the less I'm convinced that it's a desirable feature. Unless you tightly control the client, server and all intermediaries you'll spend a ton of time on operational overhead and the alternative is that you simply use unique URLs which will always work correctly after the initial implementation, which is also easier.
[1] http://www.w3.org/Protocols/rfc2616/rfc2616-sec13.html#sec13...
[2] http://www.w3.org/Protocols/rfc2616/rfc2616-sec14.html#sec14...
The correct answer is something like the video tag, with a choice, polyfills, or headers that are useful enough to do a small scale Vary on. None of which are very good solutions.
I don't think ~50% of users are running Chrome. Where are you getting your numbers from?
The real truth is that being the best technically rarely guarantees success in the market.
Apple, Microsoft and Firefox for whatever reason (politics/principle) either won't support defacto standards or have to be dragged kicking and screaming into doing so (see SPDY). And without those three you are never going to get above 45%.