75 karma · joined April 18, 2019
Personal(ish) Twitter - https://twitter.com/diqitally
It’s just naively showing the first 20 results at the moment from FTS or vector search.
Thanks for the feedback! I’ll make some edits.
You can actually search all the channels at once if you “deselect” the channel in the left! But I know that can be improved as well
The original post and the experiments were created before pgvector 0.5.1 was out, and we had not realized there was significant work to optimize index creation time in the latest pgvector release.
We reran pgvector benchmarks with pgvector 0.5.1. Now pgvector index creation is on par or 10% faster than lantern on a single core. Lantern still allows 30x faster index creation by leveraging additional cores.
Wiki Pgvector - 36m Lantern - 43m Lantern external indexing (32 CPU): 2m 15s
Sift Pgvector - 12m30s Lantern - 7m Lantern external indexing (32 CPU): 25s
The DB parameters for the above results (both Lantern and pgvector): shared_buffers=12GB maintenance_work_mem=5GB work_mem=2GB
The DB parameters for the previous results were the defaults for both Lantern and pgvector.
Benchmarking was done using psql timing and used a 32CPU/64GB RAM machine (Linode Dedicated 64).
Feel free to reach out if you need anything for benchmarks.
Would love if it could create Age of Origins, I always like watching the ads
We think our approach will still significantly outperform pgvector because it does less on your production database.
We generate the index remotely, on a compute-optimized machine, and only use your production database for index copy.
Parallel pgvector would have to use your production database resources to run the compute-intensive HNSW index creation workload.
There's some context on the operator <?> here: https://github.com/lanterndata/lantern?tab=readme-ov-file#a-...
That said, we will offer Lantern Cloud, our own hosted postgres offering (very soon. Happy to keep you in the loop. If you’re interested, please feel free to join the waitlist here: https://forms.gle/PouJxAWiSa63udJW8
With respect to recall vs QPS, we went ahead and generated this plot, hope this is helpful? http://docs.lantern.dev/graphs/recall-tps.png
You're right, 100k rows isn’t a reputable benchmark. We wanted to launch very quickly, and have benchmarking for larger datasets coming soon. Benchmarking is baked into our CI/CD, we take it very seriously!
Here’s a chart for INSERT latency (sorry about the formatting): https://docs.lantern.dev/graphs/insert.png
At the moment, we underperform Neon wrt this metric, but a better implementation is coming soon that will address this.
> Is the expectation with this (and the other) tools that I'll do a full index rebuild every X minutes/hours, or do some of them support ongoing partial updates as data is inserted and updated?
The HNSW algorithm updates the index after every insert. So all existing HNSW options (Lantern, pgvector, Neon, …) already support this.
With pgvector IVFFlat, you expect the performance to degrade over time, and you will need to re-index. This is because IVFflat’s index quality heavily depends on the centroids chosen at index creation time. HNSW does not have this limitation.
In both cases, you might want to do a full-index build to tune your hyperparameters.
We’re working on this in a few ways. One is automatic hyperparameter tuning. Another is supporting external index creation that would offload this to another server. Does this answer your question?
I don’t believe pgvector reports performance changes between releases.
At the moment, we run the benchmarking on Github CI, but we plan to move this to an external machine, since the results are unstable on Github machines. We’re planning to extend benchmarking across other repos and versions.