update: also see [2] for some primitive unsigned 64-bit integer operation benchmarks with the TFHE-rs library (winner in the sorting performance comparison of [1]). Equality at 80ms, addition and subtraction at 100ms, division at 8 seconds, etc.
[1] https://eprint.iacr.org/2026/1495.pdf Oblivious Sorting under Fully Homomorphic Encryption: A Comprehensive Survey and Performance Analysis, Omar Ahmed and Rostin Shokri and Nektarios Georgios Tsoutsos, 2026
[2] https://docs.zama.org/tfhe-rs/tfhe-rs/1.0/get-started/benchm...
Project intro talk from 2023: https://www.youtube.com/watch?v=kqDFdKUTNA4
https://github.com/google/fully-homomorphic-encryption/tree/...
Taking the above point to the extreme, a very simple yet mildly effective “homeomorphic encryption” scheme would be to run the first layer(s) of the ML model on-device, run the majority of the model in cloud, then run the remainder of the model on the device. But then you leak a lot of information that can essentially be used to get back the original text. (Usually in this type of scheme, to defend against this, the provider of cloud services doesn’t have access to the full model, it’s been used before on vision applications involving medical data)
On the other hand, each FHE step is a LOT more costly (e.g. elliptic curve exponentiation) than a vectorized BF8 multiply on GPU.
Unlike the sorting number case, it is probably same complexity. But no hardware support AND a massive slow down constant even if you were comparing apples to apples on hardware.
Disclaimer: I spent a decade working on crypto systems but I am not familiar with TFA’s research.
But softmax sampling, where you pick a single output token at the end and feed it back in to generate the next one, is branchy, so you need to do some extra encrypted computation to avoid leaking which token was sampled.
I'm not sure I've interpreted [1] correctly though, and would appreciate correction if necessary.
[1] https://arxiv.org/pdf/2604.04783 -- GPU Acceleration of TFHE-Based High-Precision Nonlinear Layers for Encrypted LLM Inference -- Guoci Chen, Xiurui Pan, Qiao Li, Bo Mao, Congming Gao, Chengying Huan, Mingzhe Zhang, Jie Zhang -- Apr 2026
Bitwise: X = A&C | B&¬C, Y = A&¬C | B&C
Arithmetic: X = A*C + B*(1-C), Y = A*(1-C) + B*C
Edit: Or to put it another way, one of the most basic things you can make in a circuit is a multiplexer, and a conditional swap is two minimum size multiplexers next to each other.
Ideally, give me a dial, to ask for encrypted intelligence when I need it. Kind of like a private chat, but with deeper privacy protections.
What real datasets exist where this would be a worthwhile trade off versus simply owning the hardware?
The numbers are so bad that underpowered local hardware would still beat it.
"We evaluate HE-LRM on UCI (health prediction) and Criteo (click prediction), achieving inference latencies of 24 seconds on UCI and 228 to 489 seconds, respectively, on a single-threaded CPU."
There don't seem to be any direct comparisons available, probably because nobody else has any reason to limit themselves to one single-threaded CPU with normal techniques, but for reference the AI seems to expect that normal times for conventional setups are in the milliseconds range, fairly comfortably, even on CPU. I didn't find a clean primary source to link to for this claim, but clicking through various things that don't cleanly state the situation it did seem plausible. So we seem to still be in the range of single-digit orders of magnitude slower, possibly as much as 5 or 6, which is to say, we're still talking the range where we need to take the log of the difference to get sensible numbers, we're not using percentages.
(To run it yourself, I basically just fed the URL from the HN link, mentioned that FHE is known to be slow, and asked if anything linked in the blog post gave concrete times.)
That said, there is a lot of ongoing work on GPU acceleration. Cf. the recent FHE-based CIFAR demo that runs in 200ms: https://sofar.belfortlabs.cloud/
Still maybe 1000x slower than cleartext, but progress!
https://belfortlabs.com/blog/belfort-partners-with-lg-on-enc...
- https://belfortlabs.com/blog/belfort-partners-with-lg-on-enc... - https://belfortlabs.com/blog/encrypted-fraud-detection-with-...
Then there’s a second tier of things that just make those wheels turn and if they do or don’t make ads revenue is nominally immaterial.
The teams doing this stuff at Google are purely for show, none of this makes it into any real products.
There’s the narrow exception of stuff like gboard, that does use privacy preserving ML/fed learning, but this stuff isn’t in the same zone.
I find it a bit embarrassing when Google publishes this stuff to be honest.
Edit to clarify my prior point: some of the technology makes it into the product, but the putative data protections do not.
Why?
Because there is always a work around, and ads legal will approve it every time.
I'm now curious. Who are you?
Thats the embarrassing part, you know “don’t be evil” went out the window years ago, yet still think you’re making a difference with your little bit shifting games. As if the moment the work you do bumped against against revenue it would hold.
If cloud can sell compute to enterprise that can only use FHE, then they will not give a fuck if the ad goons are disgruntled about it.
Ads doesn’t care until something affects them, GCP doesn’t.
and now this.
Part of changing the world involves imagining a changed-world.
That for some cases they are still all show doesn't prove everything is all show.
If Google Cloud can offer FHE-powered tools, a hospital can run Google’s AI diagnostic models on encrypted brain scans without violating privacy laws.
FHE effectively removes the primary regulatory barrier keeping enterprise customers off the cloud.
It narrows the 10^3 - 10^6 penalty to 10x - 100x.
I'm glad people fund things that are only of interest to nerds but this will never be useful.
Typical use cases are for doing biometric authentication without giving your biometric information, or sensitive queries using medical information. Apple has homomorphic encryption in image search. You can use your own photos encrypted into the cloud to search for landmarks in the image without revealing photos.
People can also coordinate and compare information without sharing sensitive data.
With wildlife management, you're dealing with health issues, like rabies outbreaks. That requires privacy. You want to preserve customer confidentially because it's often embarassing. And private property cameras and sensors can leak information about private citizens or kids in a neighborhood without appropriate social and technical protections.
There are hundreds of fields like this. Not just healthcare and policing.
We are likely to see a lot of the hardware required for this to move out into space data centers for batch jobs at least.
And along with that calls for reduced RF and light pollution like StarLink.
And that is going to be helped by a large number of angry liberal citizens who are being riled up about data centers. And that anti-tech rhetoric is already leading to violent responses and debate.
Which hurts the liberal cause for universal healthcare. Conservatives see angry liberal anti-tech actions and tarnish calls for healthcare reform and other liberal causes.
The people who are going to benefit the most from cheaper PPML in orbiting data centers are in many ways making it harder for the rest of us.
It's not a small issue, and I wish I had had the reputation to reduce the anger.
It seems unrelated to PPML. But PPML is a clever political gas pedal to get space control.
ex: A 7900XTX barely gets 0.5 TOPS of u/i64 naively w/ hip-direct, 5-10s just to bootstrap!
I needed more throughput for non-crypto i64 diff eqs so I slopped up a lib that uses RNS & CRT w/ Int8 GEMM... it's good for ~3.9 TOPS (~90% theoretical peak of the RDNA3) at prod relevant FHE sizes (2048/4096). This lowers bootstrap time to 200-500ms. It was basically free real estate lol
It isn't done yet (not worth the heat in the summer), going to finish it in the fall. Have been accumulating cloud credits to do CDNA3/4 validation in the meantime (If anyone has some to offer do let me know!)
It's neat but very dry, uses semantic contracts so you tell it what kind of mult you need and it chooses the validated best backend. If you're doing lots of smaller ops (512, 1024) it will use custom WMMA/MFMA kernels, dual issue, and grouped dispatch to land >70x over hip-direct.
But yeah, for hot path traffic it's probably going to be swamped by the input data rate. But I expected identity tables and cached lookup data will need to be a core component so duplicate checks is avoided in every way available.
Homomorphic encryption got significantly efficient with the first few iterations, but I don't really see the necessary orders of magnitude savings coming soon. You could reduce this by some partial encryption schemes (e.g., for LLMs you need a handful of basic operations) but a better alternative already exists: multi-party computation.
Source: I did research in this area in the past.
Cost-wise the only viable private compute is local compute. It's more expensive than cloud, but true private compute in the cloud is definitely pricier.
I don't see the need to do computations on the encrypted data here, which is what FHE would provide in addition to traditional encryption.
> and don't have access to the account anymore.
This would be trouble with or without FHE. Even if the backend wouldn't need to decrypt the data, the user will - so as soon as you actually want to show something in the app, you have the same key management problems as without FHE.
In order to do the data mining in order to make those kinds of claims, traditionally you'd need to have access to the data.
As you point out, encrypted-at-rest is solved. But what about when it's not at rest?
In-use and in-transit is when FHE kicks in. Sure, you could just do it locally, but then you miss out on the aggregate data mining. Not for advertisers, but because it helps women with their bodies. The compelling product claim is "we literally cannot read your period data." Not "we pinky swear not to" but "we actually really really actually can't!"
> "we literally cannot read your period data."
If the purpose is aggregated data for statistic, then surely the only per-user data they need centrally can already be aggregated (to some degree) on the device, e.g. send back only statistical-distribution variables of the personal data, for distributions over the 3-4 months? And at some point, does the service need to keep collecting data, once the model is good enough (at predicting ovulation etc)?
Another concern would be: If they are building a model, using user data, why should they own the model and thus monetise it (i.e. sell it back to its users) when users get no compensation for supplying that data in the first place.
A flow-tracking app should just stick to that, and purchase the model (for a fee) from a third party. The third party should concern itself with how to get the data without being able to leverage its position as a flow-app maintainer to trick or mislead the majority of its users into giving them free data.
I don't know much about period tracking apps, but is this really the main reason people install those apps? Wouldn't you be able to get similar results by simply monitoring (on-device) the cycle of the person who uses the app for a few months?
How do those apps work before they have millions of users?
All the warnings I've seen about period tracking apps were about unexpected data collection of the entered data. This would be pretty silly if the data collection was integral to what the user expects the app to do.
> Sure, you could just do it locally, but then you miss out on the aggregate data mining.
Ok, a bit of a technical question about FHE here: My understanding of FHE was that you have input data encrypted with some key (plus auxiliary inputs, if needed, that are not encrypted), then you do operations on that data and get a result that is (still) encrypted by that same key.
No questions there as long as you're dealing with a single key.
But the whole point of aggregation and data mining is to combine data from many different users, i.e. inputs that are encrypted by many different keys. Does that work with FHE at all? And if yes, by which key is the aggregation result encrypted?
I don't see how that would work without either "moving" data from one key to another - which would be practically equivalent to decryption - or getting a result that is simultaneously encrypted by all user keys, i.e. practically useless because no one could individually decrypt it.
> The compelling product claim is "we literally cannot read your period data." Not "we pinky swear not to" but "we actually really really actually can't!"
You could obviously read the data enough to do aggregations on it.
If you can do that for "good" purposes, what stops you to use the same aggregation algorithm for advertisers - except pinky promises again?
Would be interesting to read it (and no judgement!)