Chrome has transcended version numbers
codinghorror.com
codinghorror.com
Edit: did a bit of looking around and it seems to be planned for Oneric Ocelot
https://blueprints.launchpad.net/ubuntu/+spec/foundations-o-...
Speaking as someone who has worked on this problem for my own projects [1], I think I can answer this.
Fedora/Yum already supports downloading a binary diff between rpm packages to reduce download size, but this also requires keeping a cache of previous rpm's to run the patch against. There are multiple reasons why you can't rely on binary diffs against the files actually stored on the system, most namely for files like /etc/* that are more than likely modified since installation.
But the real problem with binary diffs is that unless you're doing what Google does to ensure that people stay up to date, the number of binaries you need to diff against grows very quickly, and there are a lot of edge cases to take care of.
For example, let's assume some package A has been released as version 1, 2, and 3. When A has a new release 4, you obviously want to build a diff against release 3, but then you also most likely need or want to build a diff against 2 and maybe even 1 to take care of people who haven't already upgraded to 3. And even if you build a diff against every single version ever released, you will still always need to provide a full version of the package as well for two cases:
1. New installations, or reinstallations, of the package.
2. When the user has cleared their package cache to save room.
And even beyond that, creating diffs involves a lot more effort and knowledge on the part of the packaging team because they not only need to know how to build those diffs, but they also need to keep track of old package versions to build those diffs against.
The end result is that you trade download bandwidth and time on part of the server and end users for a lot of effort, time, and storage space on part of the packagers and distro mirrors. For mirrors that are already encroaching on 50GB for a single release of Ubuntu and/or Fedora, adding a whole bunch of binary diff packages will most likely grow the repository size by at least 30-50%, if not more, depending on how many old versions you diff against.
The question then becomes: does this trade off actually make sense, or does it present further roadblocks for contribution from packagers and donated mirrors?
[1]: If you would like to see how I handled this sort of task, I have a Python library I wrote to handle the client side updating. I know it's not the entire piece of the puzzle because it doesn't cover generating the updates, but it might be useful for someone else. http://github.com/overwatchmod/combine
1. Diff against the previous version and store the diff.
2. Delete the previous version.
You could upgrade from any previous version by applying all the diffs in sequence, and you only need to keep one full version around. You could also discard diffs after a certain date because, as you point out, the worst case is that the full version is used instead.
I guess this would increase disk space for the mirrors, but even 100GB wouldn't be a lot of disk space, and the savings for (presumably more expensive) internet data transfer would be a lot bigger?
I didn't say you need to. Eg, with my software project, I wrote an update system that supported diff upgrading against the two latest versions, and anyone still running an older version had to download a full update.
> You could upgrade from any previous version by applying all the diffs in sequence
At that point you also need to make sure that you aren't downloading more in the process of applying a series of patches than you would need to download for a full update, which also means you need to start being aware of multiple update options, which balloons the complexity of your update code.
Well you wouldn't really need to but it would be a good sanity check on the client side.
The common case could be that people are simply keeping up with the stable edge w/o patching binaries themselves out of band---that's the case with some desktop linux variants.
The obvious way to handle this is to store full package "snapshots" on "major" version releases (probably upstream releases for packages with lots of local patching, or "one level up from the bottom" releases) and diffs in between. That is not a lot of code if you already have a sane way of managing release numbers within your package manager, which you hopefully do.
I think you can use the files stored on the system itself in many cases, at least for binaries. /etc/ and other configuration is an exception, which you could special-case. As config files are generally small files, this is no problem.
Of course you should check whether the file you are going to patch is the file you assume it is, but this is easily built-in to binary diff using a hash.
If a file doesn't match, err on the safe side and simply fetch the entire package.
You wouldn't need a diff for every combination (1->2, 1->3, 1->4, 2->3, 2->4, 3->4 in the four version case), just the three diffs 1->2, 2->3 and 3->4. Then if someone has v2 you send out the diffs for 3 and 4 and have the client apply both in order to make the updated package ready to apply. The saving in space and number of diffs stored will grow as the number of versions grows.
This will be a little less efficient in terms of bandwidth use on average when people are skipping a couple of versions, but will make little or no difference if people are upgrading in a timely manner (so are only moving in single version steps most of the time) and will save space over storing diffs between all versions.
As well as the diffs I would store checksums for each version and have the client send the checksum for the version they have just-in-case, to avoid sending a diff (sending the file package instead) if the reference file seems corrupt.
Also if you store diffs for both directions you can serve old versions of packages (in case people need to roll-back due to some unexpected incompatibility, or they are developers needing to build a test environment with older library versions) without storing every version completely. This increases the diffs per package, but not nearly as much as storing one diff between every version (for 11 versions, 1 diff per change is 10, diffs in both directions is 20, diffs between all versions totals 45 (or 90 for both directions), for 21 versions those numbers are 20, 40, 190, 380).
edit: not that I don't think this can be improved, nor that the updating software they used was any good. Just sayin'.
Of course this means that if using the multi-diff method of update distribution you would need to be careful about your selection of compression arrangement to avoid the same inefficiency (unless the saving in bandwidth for the client and the package storage servers is far more important than a bit of extra time spent on the updates client-side)..
Then it downloads the diff and the new file hash :-)
Given that for updates there is a high likelihood that the previous version is on-disk, sending some type of diff would be very beneficial to end users. Even if it is a diff for only the previous to the current version, with a full download required outside that window.
I too think this should be a much higher priority than it is for many. Fedora has had (non-default) support for this for years, but that's about it. You shouldn't worry so much about diffing against previous versions -- if you diff against the last two versions, it won't use much extra disk space, and the worst case scenario is that someone has to download the full package as a fallback, which everyone has to do now.
I think it's really a non-issue and it's not really worth talking about: Chrome just doesn't display the '1.' (or '0.' depending on your view point ^^) in front of its version number :-).
Having said that, though, I quite like Semantic Versioning[1]. The advantage it has over a single incrementing counter is that you know when API compatibility changes.
[1] I guess it can deprecate things, as long as they are still available for use.
Our online infrastructure is broken in ways we're dimly aware of, because it has always been that way. In the same way that people trying to do business demand network, electric, and roadway infrastructure that once didn't exist, we will someday demand software infrastructure with features that do not exist today.
Chief among these will be security features. If Google plays their cards correctly, they can create an ecosystem that stays ahead of the black-hat hackers. By correctly incentivizing white-hat hackers, they could expose and patch security holes fast enough to ruin the economics of the black-hats. This infrastructure will enable Google to make more money, resulting in a virtuous cycle.
If the infrastructure can be extended to the server-side, with web app frameworks that receive security updates with equal rapidity, then Google can establish a secure, smoothly running "toll road" -- an infrastructure subset relatively free from problems faced by the rest of the net. That could be worth billions.
(We'll know this strategy is winning if/when Microsoft starts doing it too. Once that happens, we'll be in a new era of computing.)
They already do it with Automatic Updates. Turn the update dial to 11 and let your machine apply them at night. I don't believe they provide binary diffs for updates, but I believe it's for logistical reasons rather than technological (e.g., title updates over XBL are surprisingly small).
Of course, MS also hasn't figured out how to update components in-place while they're being used, so expect your machine to be restarted in the morning. :-/
The part before the ellipsis is contradicted by the part afterwards. Also, Microsoft will have to get the patch to the update mechanism very quickly as well. Their record with that has also been poor. Otherwise, they will not meet a goal of keeping ahead of the black hats.
Just because your military has tanks, jet fighters, and assault rifles, it doesn't mean they're on the same level as everyone with the same equipment. There are significant organizational factors at play.
For what it's worth, this is already available in Erlang (although it was built in for different reasons, closer to getting the fluidity of web applications updates on just about any server software): two versions of the same code can live in parallel in the VM, and there are procedures for processes to update to "their" new version without having to restart anything (basically, you switch functions mid-flight and the next time an updated function is called the right way, the process just switches to the new code path).
You need follow a few procedures and may have to migrate some states, but by and large it's pretty impressive. And it could certainly be used for client-side software. The sole issue I'd see would be the updating of a main GUI window in-flight (how do you do that without closing and re-opening it?). But I doubt this one changes that much in e.g. chrome these days.
The issue I have is for the "static" chrome around the mobile parts: title bar, URL bar, that kind of stuff. There isn't much opportunity to switch that to a new process, I think. Especially if there are changes to make to the UI.
And there are just so many features like this that we need to get all the way down to the OS level before we can fully harness them, and we aren't going to get them in C. We need an upgrade of our fundamental programming primitives.
I had a call from someone who'd been using Chrome to regularly print a web page, and one day it just stopped working. The site hadn't changed, but for whatever reason the latest version of Chrome just didn't render it. And of course trying to install an older version of Chrome was quite difficult.
(In Google's case they do now have a way to disable the updates, but not all software is so good about it)
https://bugs.webkit.org/enter_bug.cgi?product=WebKit
We have a large number of tests, but sometimes we miss things. :(
At this point, I don't even bother 'updating' (read: close the browser and open it again) for up to a week or 2 after an update comes out, unless I need to close my browser for some other reason.
http://googlesystem.blogspot.com/2010/07/google-chrome-canar...
It's impressive how stable the nightly has been.
You can install Canary "side-by-side" with another channel, so you can switch back to something more stable if canary goes pear-shaped.
I guess this means for ignorant users this is good but for power-users we are having more and more control taken away from us.
Personally I disable all of Chrome's phoning home because it's impolite and does it too many times per day and I have no easy way to verify exactly it's sending all those times.
Those two plugins have 90% of people's plugin needs covered.
You can always use Firefox or some other browser instead.
Plugins are an architectural mistake. They should go away and the Chrome team is doing the right things to make it happen.
The automatic update process will also work if the webserver has write permissions on the files (which is a bad idea in the first place). But if it doesn't and you can't/won't give it FTP credentials for a user that does, you do need to go through a somewhat manual process.
Personally, I use vendor branching (http://svnbook.red-bean.com/en/1.1/ch07s05.html) in both SVN and Git in cases like this. I don't have to rely on the developers to generate a patch: I get all of the changes pulled directly from my local repository.
1. Updates should not only be applied in sequence.
It is better to produce a binary diff between any two versions, and apply only that (one) binary diff. The reason for this isn't efficiency, but semantics. Updates not only fix things, but break things. Meaning, updates corrupt application state (data), both in-memory and on-disk. It can be disastrous to apply an intermediate update that removes state, only to realize that a future version reversed the semantics and needs to use that state (which was available, but is now gone).
Peserving backward compatibility is important, which means the ability to skip some version updates is necessary. To the extent possible, reversing updates is important too.
2. The ideal update system should apply updates live, not offline.
With a model that accounts for updating the entire state of an application, updating live is possible. The reason most updates are not applied live yet is that the model is not descriptive enough to change the entire state of the running application.
Notable state that should be updated, but often isn't, is continuations and the stack. This is why GUI applications need to be shut down to update.
Scheme's call/cc (call-with-current-continuation) solved making changes to continuations and stack state decades ago better than Erlang. Erlang cannot force stacks unroll or continue from arbitrary points.
3. Updates must be produced with source code and programmer input.
Updates should not be produced with binaries as input.
The reason is the need to account for application semantics, which binaries do not expose in the detail source code does. Although automated, sophisticated semantic-diffing based on control-flow can be developed, it is sometimes inconclusive whether an update will break things.
4. It is necessary for programmers to provide live update guidance.
In the cases where producing provably safe dynamic updates is not possible, it is input from the programmer that can clear any conservatism of the safety certification process.
Tools are needed for programmers to reason about the semantic safety of their live updates, integrated in the development process. Including tools that help transform application state between versions.
I don't think it ever displayed that dialog on OSX.
Edit: wording