Glad to see more work on pgvector but why test on such small datasets on a large memory machine? The big ann datasets have 1B points and are much more interesting/representative of current embedding use cases (eg from dual encoder models).
I’m also curious if there is a way to not store everything in memory for pgvector. Is that possible?
Lastly, what is the parallelism story? Is it just using a thread pool under the hood? OpenMP?
Understanding if pgvector plans to support point insertions and deletions is also important in practice.