My friends work on the SCaNN team!
ScaNN doesn't use a second ML model, it's just an efficient way to store a bag of vectors. But if the software is too heavyweight for you, you don't necessarily need all of the vector quantized multi-level trees and bit twiddling tricks, you can implement 20% of the work for 80% of the speed gain.
Here are a few super simple approaches:
- Random projection: instead of doing the search in 1024-dimensional space, randomly project your vectors down to 16 dimensions and do the search in that space instead. Objects that are far away in this smaller subspace are at least that far in the original space, so you can use this heuristic to prune most of the dataset away; then, you can rank the closest items using the full nearest-neighbor search to get exact results.
Dead simple, lossless heuristic, speedup factor is (new dimensionality) / (old dimensionality).
- Locality-sensitive hashing: in addition to storing the vector representation, store its hash. This can be quite simple, e.g. random projection LSH converts the vector into a series of bits, according to which side of a random hyperplane that vector falls on; see https://www.pinecone.io/learn/locality-sensitive-hashing-ran... Unfortunately, this will be "lossy" - some images will be missed if they fall into different hash bins, and far images may be hashed to the same value if the region w/ equivalent hash is thin/narrow/oddly shaped.
More complicated, lossy results, can search unlimited images in constant time.
- Multi-tree lookup: break the space down into a KD-tree and search that instead. Complicated, and it stops working in fairly high dimensions (certainly don't use this on 64-dimensional vectors or higher)
ScaNN integrates most of these techniques, plus a few more arcane approaches that depend on processor-tuned heuristics. The gory details are available in the team's ICML2020 paper, see http://proceedings.mlr.press/v119/guo20h/guo20h.pdf