Why SpiderOak doesn't de-duplicate data across users
spideroak.com
spideroak.com
First, you can do variable block-based de-duplication, which is how major storage vendors do it - not per-file, which doesn't really buy you much.
Leveraging this in the SAN firmware also prevents their ridiculous file transfer 'vulnerability' (which only exists due the way they wrote their software) - all of your files copy over the network to the storage system. Once they're on the storage system, at some later time and asynchronously, the storage system runs a dedupe on the blocks, and winnows down its storage. Think of it like transparently compressing on the storage side, only, hopefully, less intense on I/O.
Finally, they could simply encrypt everything and then they can't answer subpoenas about who has knowledge of what - it was all encrypted immediately after uploading, and no log is kept. If the data doesn't exist, you can't be forced to give it up...
I disagree that requiring end-users to effectively operate their own encryption software is a robust path to privacy.
The blog doesn't mention whether or not SpiderOak uses block-level deduplication. Maybe it's part of their storage infrastructure, maybe isn't. But all that client-side encryption would severely reduce the number of duplicate blocks even if everyone uploaded the same file.
So you pick a hash function whose space is so large that the risk of collision is less than the risk of any other possible reason for accidental mis-identification (like all the file's bytes spontaneously switching to the collided file's bytes). You can have databases with trillions of objects with less than a one in a million chance of collision in a 256-bit space.
Realistically, there are a lot of things that are better to worry about than a one-in-a-million chance of losing a file. And if you really need so many objects that that's not enough, just increase the size of the hash space.
True, but this does not happen in practice. Using SHA-1 as an example hash function, you'd need about 10^24 different files before you would expect a 50% chance of a collision. You are not going to come anywhere near this limit, 10^24 different files have not been created during the entire span of human history.
Regrettably, we don't actually have any truly ideal hash functions yet.
Of course, turning a collision into an exploit in this case would be challenging. (Preimage attacks ⊂ collision attacks.)
This explains why adding the ubuntu netbook iso image to the Dropbox folder, synced in milliseconds.
My friend was going a little nuts wondering if Dropbox choked.