Mega, the new MegaUpload, to launch on January 19
kim.com
kim.com
It seems a good idea to me, maybe because I thought of it before :-) I'm not sure how well they will monetize it. I've read that economies come from detecting duplicates and archiving contents just once. But from the customer POV, that would be a strong feature.
Is there a way to store files encrypted with different data and detect they are the same file ? If there was, wouldn't this allow LEO to verify this and issue DCMA notices ?
I understand is possible to do operations on encrypted data (which results in an encrypted result) all without having the key to the data (or the result), and maybe there's a way to do this to allow the duplication.
If there is it could (but shouldn't ?) be used.
The issue here, to me at least, is that someone malicious that has exactly the same data could encrypt it and then verify the above.
It's also highly experimental at this point, from what I remember.
[1] As a trivial example, let's say you give me a very large number, and your encryption scheme is to add some number, n, of zero bits at the end. Only you know how many bits you're adding - n is your private key. Regardless of what you pick for n, I can multiply your number by 2 (i.e. bit-shift it) and give the result back to you, which you would then be able to decrypt to the result of the calculation.
This works for any value of n, so I can't tell if two original numbers are the same by inspecting the ciphertexts.
The keyword is "convergent" encryption. We used something like this at Iron Mountain Digital many years ago (they still do, AFAIK), and it is used in BitCasa today.
You should read the papers, but essentially the concept can be boiled down to encrypting the plaintext with a hash of the plaintext.
Since there is no way to derive the hash of a plaintext from an encrypted block, there is no way to hack the key other than regular old brute force. But if the same data is uploaded twice, the same hash is computed, and thus the same encryption is used, and thus the encrypted cipher text is identical.
The encryption keys can be stored separately from the cipher text. In particular, the user who uploaded the data would store the hashes (this would already happen in most backup applications anyway). Then, for retrieval, they give the hash and the block location to the server, who is now able to decrypt it. By stealing the server, you gain zero access to plaintext data.
Very cool stuff :)
But your 'random salt' idea suffers from the attacker just generating all possible encryptions of the plaintext due to the small number possibilities. The "convergence key" is solely a security-parameter-length key that you can pass around to your friends so that your files will dedupe with theirs while not being susceptible to confirmation attacks by others.
I haven't tested cyphertite, but I've been meaning too. I mean, Ryan McBride is involved in the project as well as other OpenBSD devs. I'm hoping it has the same level of polish as OpenBSD.
That's a nice catch-22, as you need contents of the file to obtain the key to decrypt it.
Deduplication could be even more effective if the file was first split into variable-size content-dependent blocks using rolling hash (like rsync does) and then each block was encrypted this way separately (this way same MP3s with different ID Tags would still be mostly deduplicated).
Of course the more you make deduplication easier the more you indirectly disclose about contents of the file, so this is a security/privacy trade-off.
For each piece of content (e.g. a photo or video), generate a ~32-byte random string, and symmetrically encrypt the content using that random string as the key. Then encrypt the random string N times using each of your N friends' public keys and give them the result. That way you control access per-item, people you unfriend can't decrypt content you post after the unfriending, and the content itself only has to get encrypted and stored once.
Is that how your thing works? Was it successful from a technical perspective?
Besides, it'd add more complexity to an already sensitive algorithm. As the Zen of Python says, special cases aren't special enough to break the rules.
Was privacy really ever their main concern ?
I expect it will be with the authorities breathing so hard down their necks, and the questionable legality of the files users will likely upload.
>Unfortunately, we can't work with hosting companies based in the United States. Safe harbour for service providers via the Digital Millennium Copyright Act has been undermined by the Department of Justice with its novel criminal prosecution of Megaupload. It is not safe for cloud storage sites or any business allowing user-generated content to be hosted on servers in the United States or on domains like .com / .net. The US government is frequently seizing domains without offering service providers a hearing or due process.
It's a pretext for hiding what content is allowable for download to protect them from liability.
Ps. You can access it here too, http://mega.co.nz/
Their me.ga domain was taken away from them.
The real bone with megaupload was, authorities held megaupload responsible for what it hosted - and Megaupload could not deny what they had on their servers (copyrighted stuff) - so the responsibility and liability lied with megaupload and caused its downfall.
With mega.com - the game is, mega.com will claim we don't know what is on our servers (since its encrypted at browser) so we can't be held liable for it - and this will stick!
I read somewhere comment that - its doesn't matter how strong the encryption is for mega.com users - all mega.com need it as a shield from legal troubles.
Clever!
In the previous model of receiving data then possible encoding it, they have full access to the raw data uploaded and are responsible for policing the legality of the files being uploaded.
What does Instra provide exactly ? I used them for domains but I don't really know if they also provide some good hosting infrastructure.
http://www.nbr.co.nz/article/nz-company-named-key-mega-partn...