I've seen a few FHE posts roll across the front page recently and they all make me think of Vaultree because they sound like they've got it sorted.
I've seen a few FHE posts roll across the front page recently and they all make me think of Vaultree because they sound like they've got it sorted.
For those interested in the technical specifics, Vaultree has developed a comprehensive approach to searchable encryption, detailed in our patent (EP4000213A1). This method enables efficient and secure search operations on encrypted data, ensuring that sensitive information remains protected without sacrificing usability. You can explore the full details of our patent here: https://patents.google.com/patent/EP4000213A1/en?q=(Vaultree...
Additionally, our work on Fully Homomorphic Encryption (FHE) represents a significant leap forward in the field. We've published our FHE scheme in the IACR ePrint archive, where it is accessible for review and further academic scrutiny. You can find the publications here: https://eprint.iacr.org/2024/1105 and https://eprint.iacr.org/2024/1622
Moreover, we are actively working on encrypted Machine Learning (ML) implementations. To contribute to the broader community, we've open-sourced our VENumML library, which is based on Vaultree's FHE. This library aims to enable secure and private ML operations on encrypted data, pushing the boundaries of what is possible in privacy-preserving technologies.
Vaultree's commitment to innovation and responsible encryption practices ensures that we not only keep data secure but also advance the field with solutions that are both effective and practical for real-world applications.
Encryption as a service is non-sencical. If the provider has the key, and does the encryption and decryption, then who are you protecting the data from[1]? What magical malicious person are you imagining that would somehow be able to get their hands on the encrypted data without also getting the key?
[1] this is very different from FHE where the provider recieves the data already encrypted, and at no point has access to the decryption key.
Edit: their website claims "Data is never decrypted" but then claims they decrypt it before returning it. So its confusing what they are actually doing - but i am 99% sure they are selling bullshit.
If you're using FHE to encrypt a text search, then you'll generate a match/no-match boolean for each character. The server won't know which is which. Then you'd probably OR large blocks of these together, to give you a match/no-match boolean for each segment of text. Then you return all the booleans to the client, who decrypts them.
Are you saying something along the lines of - they split the data in to tokens, deterministically (no iv) encrypt each token, and then do equality comparisons on the encrypted tokens?
Maybe, and it would explain why they say you can chose a cipher and then list a bunch of standard symmetric ciphers. However such schemes usually leak too much in practise (even at the granuality of whole words).
More importantly, it really doesn't matter. They have both the key and the encrypted data. If someone hacks their system, the best encryption in the world won't help if the attacker steals both.
If that's the case it simply moves "the need to trust the DB host" to "the need to trust the encrypt/decrypt intermediary" (literally a MITM, LOL)
> Vaultree's proprietary encryption breakthroughs are in various encryption technologies traditionally limited to niche use cases. We finally enable users to process entirely encrypted data with Fully Homomorphic and Searchable Encryption (FHSE) and other technologies in the field. Explaining what they are would take all day, but here's a one-liner: FHSE enables data processing to be run directly on encrypted data in the same way as on plain text data.
Absolute bullshit.
> You choose the encryption standard in use for the database, from AES, DES, 3DES, Blowfish, Twofish, Skipjack, and more.
That seems very wrong ((as far as I know) those standards are not in any way designed in such a way as to permit operations on their cyphertexts).
> Vaultree has achieved major breakthroughs in several encryption technologies, allowing organisations to process fully encrypted data at near plaintext speed and keep their data safe even in case of a leak.
That seems even wronger (security via obscurity at best).
Those standards are meaningless by themselves without specifying a mode (e.g. GCM, CTR, CBC, ECB, etc). A thing people sometimes try to do with them (no idea if this is what vaulttree is doing) is use some less secure mode that is determistic and do equality matching (this is almost always a bad idea and usually leaks way more than you would naively assume). For example, if you use ECB mode you can search as long as you are searching along block boundries.
https://www.microsoft.com/en-us/research/wp-content/uploads/... is an interesting paper about this sort of thing.
I could see a use cases in either defense in depth and/or storing data in the cloud while having your keys somewhere else.
I could be wrong, but it smells like snake oil to me.
For example, they give an example of running queries against data that is never decrypted. I'm very curious as to how they do this. I've used blind indexes [1] to solve the "encrypted data searchability problem" in the past, but with blind indexes you're still left with the fact that you can only do exact matches - you can't sort the data or use less than/greater than queries.
With true FHE you should be able to sort results, but my understanding is that it's several orders of magnitude slower than plaintext searching, so I'm very curious as to what Vaultree is actually doing.
1. https://medium.com/@joshuakelly/blind-indexes-in-3-minutes-m...
If you skip the security requirements, applying rot13 twice is a fully homomorphic scheme that achieves plaintext speeds ;)
Our insight: if you focus on a specific problem, you can apply FHE much more efficiently and end up with practical speeds. (General-purpose FHE is still probably a ways off).
We are particularly interested in the problem of private information retrieval - fetching items from a large database, without revealing anything about your query to the server. Our server (open source! [1]) supports private queries against gigabytes of data in under a second.
That's a performance level that enables cool apps today. If you have any ideas for using FHE, do try out our SDK [1].
Techniques I'm seeing in the Pappas et al. paper mentioned in the history section of [1] to do more complex queries seems pretty cool, and I imagine the performance has been improved a bit in more recent work.
[1] https://en.wikipedia.org/wiki/Searchable_symmetric_encryptio...
Up can't do a range on encrypted data. If you encrypt 5 and encrypt 10, how do you expect to compare the encrypted results to see which is greater?
If all you do is key value lookup then sure. But SQL is much richer than that.
See this discussion about how to achieve that: https://news.ycombinator.com/item?id=31668814