Short answer is most of these databases uses some type of precomputation to make doing approximate nearest neighbors faster. HNSW[0], FAISS[1], SCANN[2] etc are then all methods of doing approximate nearest neighbors but make use of different techniques to speed up that approximation. For your use case it will likely result in a speed up.
[0] https://www.pinecone.io/learn/hnsw/ [1] https://engineering.fb.com/2017/03/29/data-infrastructure/fa... [2]https://ai.googleblog.com/2020/07/announcing-scann-efficient...
All of this assumes you're okay with a bit of imprecision - vector search with modern indexes is inherently probabilistic, e.g. your recall may not be 100%, but it will be close. Using a flat indexing strategy is still an option, but you lose a lot of the speedup that comes with a vector database.
Q[i,:] /= norm(Q[i,:])
A[k,:] /= norm(A[k,:])
So now with that preprocessing the cosine similarity of a given row i of Q and k of A is: cossim(i,k) = dot(Q[i,:], A[k,:])
If you multiply QxA.T (10,000 x 128)x(128, 1M) you get a result matrix (10,000 x 1M) with all the cosine similarity values for each combination of query and vector.If you make a pass across each column with a priority queue, you can find the top-n cosine similarity values in time O(1,000,000xn).
Now you could store the resulting matrix, but Q is going to change for each call, and we really only care about the top-n values for each query, so storing it wouldn't really accomplish anything.
Edited: fixed lots of typos
pastel-mature-herring~> Could this matrix be compressed to binary form for storage in a binary index?
angelic-quokka|> It is possible to compress the matrix to binary form for storage in a binary index, but this would likely decrease the accuracy of the cosine similarity values.
This site has some nice information on how ANN performs for vector search.
https://www.pinecone.io/docs/api/operation/query/
So maybe it wouldn't work for you use case.