On-disk HNSW index for Postgres with pg_embedding
neon.tech
neon.tech
What the work done by Neon, the pgvector team, Supabase and others points to is that "speed" isn't the only factor in vector database selection. Developer experience and existing infrastructure investment are too.
HNSW-IF is an excellent extension to HNSW (that the vespa team has made easy to implement) that takes advantage of the speed/recall of HNSW in combination with the disk scalability of inverted indices - it is a hybrid approach.
https://unum-cloud.github.io/usearch/
Which was also recently on HN: https://news.ycombinator.com/item?id=36942993
Does the the initial construction (creation) of the index need to fit into RAM?
Also, should one expect the reading via Postgres's buffer cache to have better access patterns / need less memory / have other benefit vs having a non-Postgres HNSW index that's mmapped from disk?
I'm interested in how many vectors are indexed/how large the index is that corresponds to the latency chart? If we have an in-memory HNSW index of 10M vectors at ~20GB (512 dim), say, what are the RAM requirements when using the disk-based version?
What's the difference between pg_embedding, pg_vector, and tsvector? Are they compariable/interchangable? And how do you know which one to pick?
My understanding is pg_vector has poorer performance compared to some dedicated vector databases, does pg_embedding perform better?
Sorry if these are silly questions.
In this case "hnsw" is the name of the algorithm. If you want to know more, just search "hnsw" and "ivfflat" to understand the differences.
tsvector is something else; it's not a numeric vector index, it's a different kind of data structure that stores the actual lexemes (keywords) with positions and weights to facilitate full text search (rather than using a text embedding model.)
However pg_embedding serves the index out of disk, whereas most vector databases opt to serve the index out of memory, or delegate to mmap'ed files. HNSW is a graph-based algorithm, thus access patterns are random and this does not lend itself to disk based access. I'd expect pg_embedding to be slower than memory-resident indices due to this fact. Also in general with a postgres index, my concern would be scalability and resource isolation. It's convenient to colocate these things but ANN indices have very different CPU/Memory/Disk usage patterns than what you may need for just your relational data.
For example
Chroma -> Serves HNSW out of memory, persists to disk with a WAL.
Weviate -> Serves HNSW out of memory, writes HNSW graph search to WAL and uses that for durability.
Milvus -> Serves HNSW out of memory, supports partial mmap, also supports another algorithm called DiskANN which is optimized for SSD. It uses cloud storage with a shared-everything architecture for durability (a lot of nuance here.)
QDrant -> Has a WAL, supports memory and mmap'ed indices.
Many teams can also get away with good performance vs the fastest performance, given smaller index sizes and the other tradeoffs I mentioned.
That said, I can imagine the pgvector folks precaching the new HNSW index support they're working on, as they do with their IVFFLAT index.
* edited for the grammar gremlins
The commenter - thewataccount - asked about performance and I shared my intuition.
Also the cache access patterns will vary between these implementations and is worth considering.