Why Sync Is So Difficult
gigaom.com
gigaom.com
Yes, sync is hard. Where I think things fail is mainly around user experience. You just have to make smart decisions and let users recover if they see something they don't expect. Dropbox does this quite well.
I learned this when I was at Microsoft on the ActiveSync team -- we actually were trounced by RIM not because our sync was worse (it was probably superior) but because we initially made the mistake of OVER-reporting status (even minor conflicts that normally you would want to just ignore.)
Sync should be a silent, no/very little UI experience -- a utility that just works in the background. Any attempt to make it more than that will cause the product to fail miserably.
There is no way to automate this conflict resolution for binary files (if the same file is modified in two places). You can simply pick the latest modified version, but this is less than optimal (the other modifications will be lost then). Even dropbox isn't totally silent, in a conflict such as this, their server keeps revisions (so you can pick and choose the correct one, making it not silent).
Silent Sync (in a automated no-user-interaction way) has , is, and always be a dream (at least for the foreseeable future).
There are a ton of other good features in Notes (and a few bad ones too).
I implemented my own 2-way syncing for ShoveBox and its new iPhone app (wonderwarp.com/shovebox).
I was very thorough in the way I did it, but there are still small issues that I wasn't able to resolve by release time.
I'm going to write a blog post on this soon.
Yes, of course. That was exactly the point I was trying to make: it's a fundamentally hard problem.
> Git and it's way of handling trees of old commits is a decent place to start. It clearly isn't a final solution, but building on top of it seems like a worthwhile direction, at least to me.
That depends on what problem you want to solve. If you want to solve the general data-storage-in-the-cloud problem, then Git is fundamentally flawed because 1) one of the inescapable aspects of the problem is that the solution depends on the semantics of the data and 2) Git by design knows nothing about the semantics of the data.
I don't think using git itself directly would solve many of the hairy platform issues, because they are really outside the scope of what git itself tries to do.
There's also the complication that if you want a zero-knowledge approach to privacy (where server admins can't read your data OR your file/folder names) some new data structures need to be invented.
My thought would be to store some meta-data on devices about the current version of the file, and just check the remote server for a newer version before it's opened.
SpiderOak implements a different approach, having initially built a comprehensive journaling backup. Sync happens as a result of logically combining the journal entries from all available end points. There's no "event replay." The final state for a folder is calculated based based on the user's likely intent from the totality of all actions taken in each folder over time. The set of actions to perform locally on any device is the diff between the calculated end state and the local state.
There are still some good points here. The cross platfrom issues mentioned are subtle. (and SugarSync doesn't even support Linux.) For instance, ":" is a valid character in Mac/Linux filenames but not Windows. And the case sensitivity/insensitivity can create conflicts where they wouldn't otherwise exist.
For character encoding, Windows is actually the easiest with Unicode natively stored. Mac and most Linux distros use UTF_8 but there's nothing stopping users from dumping a bunch of filenames with arbitrary heterogeneous encodings all in the same folder.
You're right about the fact that there wasn't historical versions. But that was a year back. They're here now, have you not seen them?
Unicode encoding is very subtle. The difficulty is that there are several different - but equivalent - unicode encodings for the same strings, and the different filesystems use different normalizations to make sure that they can compare their strings byte-by-byte.
I would add that file locking issues is also a huge problem even when it comes to a simple conflict resolution.
Throw in case sensitivity issues, among others and yeah, sync is difficult.
Why Sync Is So Difficult ?
2-phase commit uses all or nothing and asynchronous replication does not use it.