The most private AI is the one running on my own hardware, not in some giant data center.
The most private AI is the one running on my own hardware, not in some giant data center.
I want that too, but you gotta ask yourself the question how efficient that is compared to running it in a datacenter shared with everybody else.
As an aside: The computation might also not be the same e.g. ever-changing hidden pre-prompts, security/safety checks blocking or degrading responses, unavoidable verbosity to simple questions, watermarking, collection of prompt data to build user profiles for the purpose of advertising - and we haven't even seen in-response adverts, or sponsor-biased responses yet, but no doubt it's coming.
Using regular encryption and secure enclaves, there are already providers that are roughly 2x the cost of normal providers. For example, https://tinfoil.sh/
Using this other encryption, the provider has neither need nor capability to decrypt it on their end, so the user gets extra security.
the main difference is where the guarantee comes from. for FHE, it comes from math, which we trust. for secure enclave, the guarantee comes from Intel/AMD's promise that their hardware is bugless/backdoorless, and that your adversary cannot directly inspect bits in the hardware
- The output can reveal information to the provider, which homomorphic encryption would have protected
- Inference is running on GPUs - so its moreso nvidia than amd/intel, but this is just a nit
So homomorphic encryption exists so the user doesn't need to do work to figure out if the provider could be adversarial.
That's why they provide cryptographic attestation that the open model they're running is exactly what they advertise without any modifications.
That combined with GPU confidential compute should protect your LLM prompt and output.
E2E encryption (including homomorphic encryption) have the nice property that there are much fewer ways for things to fail.
(Tangentially, attestation is basically trying to ensure that faults are obvious, but that doesn’t reduce the probability of the faults in the first place).
1. A ZDR clause is "trust me bro". You have zero way of verifying their pinky-promise.
2. A ZDR clause is still subject to the old-classic "government, court or administrative order" catch-all clause. :)
3. "Even with ZDR enabled, Anthropic may retain data where required by law or to address Usage Policy violations. If a session is flagged for a policy violation, Anthropic may retain the associated inputs and outputs for up to 2 years, consistent with Anthropic’s standard ZDR policy." (I quoted Anthropic, I'm sure all the others have similar).
Heh yes absolutely, but there is some nuance.
Secure Enclave still requires you to trust the operator and also trust that it’s configured properly, supply chain is secure, etc.
The beauty of FHE is that it doesn’t rely on the compute being secure. All you need to secure are things you already have control over as a client.
I agree with you it’s still way too slow to be generally useful. (By general, I mean practical for arbitrary computation — you can relax the requirement and have fast homomorphic encryption if you only do specific kinds of operations).
> Fourth, there is a bandwidth concern. FHE encryption schemes generally increase the size of the data being encrypted, and the user must send the server a special set of encryption keys to enable the computation, which are relatively large as well. The special keys need only be generated and sent once and can be used for all future computations, but they can easily be gigabytes in size. In one example FHE scheme with lightweight keys, a ciphertext encrypting a single integer is on the order of 25 KB, and the special keys are about 0.5 GB. In others, 16,000 or more integers are packed into a single ciphertext of similar size, but the keys can be 10s of GiBs.
Sibling comment estimates lower bounds of current research at minimun 10^6 overhead which sounds more realistic.
There is no reason to believe it should be lower than that - or even that low. Or do you have access to research claiming such achievements?
I was very much surprised and asked. give me demerits for the way of asking.
but the question stays: how come an encryption scheme inflates data by this order of magnitude and needs GB sized keys?
where can I learn about this? not the nutty gritty details proofs and all but an overview. assume I did my CS masters in the 1990s and worked as SW eng ever since.
NVM, I asked Gemini