Pgvector 0.5.0 Feature Highlights and how tos
jkatz05.com
jkatz05.com
SELECT * FROM A, LATERAL (SELECT * FROM B ORDER BY a.embedding <-> b.embedding LIMIT 2) AS closest_in_b ORDER BY A.id;
It’s going to have so-so performance. So, don’t hose your production server. Create a vector index on b.embedding. HNSW will be nice for this use case.
If anyone is into embeddings, check out Instructor Large/XL. It's quite good and super fast using L4s. Haven't quite figured out the instructions bits yet, but got it clustering things today and that was cool.
pgvector_hnsw outperforms pg_embedding using a query per sec / recall measure across various common embedding widths and a range of dataset sizes.