SW-delta: an incremental cache for the web
github.com
github.com
Cloudflare has a similar solution called Railgun[2] for updating dynamically generated pages. "reddit.com changes by about 2.15% over five minutes and 3.16% over an hour. The New York Times home page changes by about 0.6% over five minutes and 3% over an hour. BBC News changes by about 0.4% over five minutes and 2% over an hour."
[1]: https://tools.ietf.org/html/rfc3229 [2]: https://blog.cloudflare.com/efficiently-compressing-dynamica...
It generated a 512 byte delta string, instead of downloading the new version of 83KB, 512 bytes seems like a pretty significant optimisation. Also, 2300 bytes for 2.2.0 -> 2.2.4.
I haven't seen a lot of great service worker uses so far but this seems plausible. Good job.
Delta if you wonder what it looks like: https://gist.github.com/eknkc/fb27cfaee871a007c3cabfda5df03a...
I wonder how the delta size compares to something like rsync or bsdiff?
I looked into bsdiff, and apparently it internally uses bzip2 for compression of several components independently. As a quick hack, I modified it to output all those pieces uncompressed (producing a large file of mostly 0s), and then tried compressing the result with various compressors, both to test other compressors, and to compress the entire file as one unit rather than as separate components.
The result ("ubsdiff" is bsdiff without compression):
85659 jquery-2.2.3.min.js
85578 jquery-2.2.4.min.js
376 jquery.bsdiff
85994 jquery.ubsdiff
264 jquery.ubsdiff.brotli
252 jquery.ubsdiff.brotli9
309 jquery.ubsdiff.bz2
403 jquery.ubsdiff.gz
360 jquery.ubsdiff.xz
So, 376 bytes for unmodified bsdiff, 309 bytes by compressing the whole uncompressed bsdiff file with bzip2, 264 bytes by compressing the whole uncompressed bsdiff file with brotli, and (strangely) 252 bytes with quality 9 brotli (the default is 11). 85578 jquery-2.2.4.min.js
86351 jquery-3.1.0.min.js
8663 jquery.bsdiff
101359 jquery.ubsdiff
8380 jquery.ubsdiff.brotli
9209 jquery.ubsdiff.brotli9
8853 jquery.ubsdiff.bz2
9550 jquery.ubsdiff.gz
8360 jquery.ubsdiff.xz
In this case, bzip2 of the whole file did noticeably worse than bsdiff's compression of three separate components. brotli of the whole file still won, though. Which made me wonder if brotli of the individual components would do better than brotli of the whole file. Turns out it does: 8006 bytes.OTOH, the entire jQuery 2.2.4 is just 26k brotli compressed, 29k gzipped.
I've heard that Google's Inbox (and likely other websites) uses this technique, although the implementation AFAIU is not open source.
...actually, I think their's is different, in that it doesn't depend on service workers; AFAIU the approach is:
if you move from js-lib-v1.js to js-lib-v2.js, they'll go ahead and source js-lib-v1.js in the browser, and then also load js-lib-v1-to-v2.js, which is a server-side generated JS file that redeclares/redefines only the JS functions/modules/whatever that have changed from v1 to v2.
So, I believe their approach is much more intricate, because I believe it diffs the JS at a semantic level to generate the "patch the already-loaded JS by doing another JS load", vs. AFAICT your approach of just doing a textual diff.
Assuming my assumptions about both approaches are right, I definitely prefer yours in terms of simplicity; albeit the Google approach is (or was?) necessary to benefit most users.
There seems to be a trend of making un-opinionated software that acts as more of a sandbox for down-stream developers. I think that's great because we are now seeing the great solutions coming out of that sandbox. But at what point do we start to have a general consensus and implement some of these great solutions natively to remove the resource overhead?
The big push in browser standards is the Extensible Web Manifesto[1]. The idea being that we should first give web developers powerful and general purpose primitives to explore and build solutions with, because there are many more web devs than there are browser implementors, and the standardization process is slow. We can then explore standardizing the most successful and useful results.
This project is using service worker, one of the canonical extensible web APIs as it gives sites tremendous control over their network use.
[1] https://www.w3.org/community/nextweb/2013/06/11/the-extensib...
JQuery functionality is for the largest part integrated in modern browsers, JQuery is living on because of momentum, it's ecosystem and the desire to support older browsers.
React meanwhile is popular in some places, but hasn't reached nearly the age and ubiquity that would warrant a native browser implementation.
Edit: Confused Shadow DOM and Virtual DOM.
Perhaps you're thinking Virtual DOM, which is a technique to apply only diff-ed state changes to the real DOM, which isn't actually standardized (yet), although alternate implementations (outside of React) exist?
[1] https://facebook.github.io/react/tips/inline-styles.html
Because there are strong standards on a platform like iOS, you see almost no competition in many parts of the stack, and so that platform is limited by Apple's developer resources and constrained by the necessity of designing the architecture for the lowest common denominator.
The general philosophy in web standards groups has been to start by providing the simplest possible API that will give developers access to the functionality they need, allow developers to build frameworks on top of that. And then as common use cases emerge, back port the most widely used features into the API.
The Service Worker saved a complete HTML rendering of the current state of the React app, which was then served on reload, so that the correct view showed up even before any JavaScript was loaded.
Then in the next "requestIdleCallback", the React app was initialized with store data that was also cached.
It only works on Firefox and Chrome, so it's great for performance improvements or if you control the client's browser setup.
Other than that, try looking at your browser's debug pane for service workers and you'll see if any site has installed them. In my experience there are quite a few.
ProductHunt uses them to send notifications when you're off the site, which is a bit annoying, but maybe I said yes to it at some point...
The sw-delta-client rewrites the URL, and the altered request is sent to the server. The server responds -- let's assume with a cacheable respone -- and the browser's cache gets updated.
Next time, we request the same URI but it's already cached and not stale, so it can be served right away from cache without having to go to the server. Does sw-delta-client intercept such a get-from-cache request? What if the cached entry is stale and needs to be revalidated by making a conditional GET to the server. Does it intercept revalidation requests?
Depending on some of these answers, the addition of query strings to the browser-perceived URI may influence whether browser caching is performed (properly, or at all). See: [1] http://stackoverflow.com/questions/24354119/ [2] http://stackoverflow.com/questions/3131518/ [3] https://support.cloudflare.com/hc/en-us/articles/200168256-W... [4] http://stackoverflow.com/questions/23603023/