1. a SIMD-optimized form of product quantization (PQ), where code to distance lookup can be performed in SIMD registers
2. anisotropic quantization to bias the database towards returning better maximum inner product search (MIPS) candidates, versus usual quantization (such as PQ) that aims to minimize compressed vector reconstruction error. In MIPS it is much more likely that the query data set may be of a completely different distribution than the database vectors.
If your application needs L2 lookup rather than MIPS which does not admit a metric, then only the PQ part is relevant. For cosine similarity, you can get that by normalizing all vectors to the surface of a hypersphere, in which case it has the same order as L2 (see "L2 normalized Euclidean distance" on https://en.wikipedia.org/wiki/Cosine_similarity ).
SCaNN is implemented in Faiss CPU, on the GPU the fast PQ part is less relevant due to the greater register set size and throughput (rather than latency) optimized nature of the hardware, but the GPU version is more geared towards batch lookup in any case.
https://github.com/facebookresearch/faiss/wiki/Fast-accumula...
https://github.com/facebookresearch/faiss/wiki/Indexing-1M-v...
We have not found the anisotropic quantization part to be that useful for large datasets, but results may vary. Graph-based ANN techniques tend to be better in many cases for these small datasets than IVF based strategies.
(I'm the author of GPU Faiss)