HNHacker News
TopNewBestAskShowJobs

VoVAllen

231 karma · joined August 3, 2022

submissionscomments
VoVAllen··on Photobox – Free Open Source Google Photos Clone
immich did this perfectly. see https://www.reddit.com/r/immich/comments/1bdi3dz/immich_is_a...
VoVAllen··on Ilya Sutskever to leave OpenAI
Why now?
VoVAllen··on AI Infrastructure Landscape
Better to have an AI agent which can add the company by one-click with name (even better if it can update with the latest news)
VoVAllen··on 90x Faster Than Pgvector – Lantern's HNSW Index Creation Time
Sift 1M is too small to make meaningful comparisons. Storing 96 floats * 1M only takes up 800Mb of memory.
VoVAllen··on Investing in new vector database development vs enhancing existing databases
It's kind of a tradeoff. Performance is just one factor when choosing the vector database. In pgvecto.rs https://github.com/tensorchord/pgvecto.rs, we store the index separately from PostgreSQL's internal storage, unlike pgvector's approach. This enable us to get multi-threaded indexing, async indexing without blocking the insertion, and faster search speed comparing to pgvector.

I don't see any fundamental reason why the index in Postgres would be slower than a specialized vector database. The query pattern of the vector database is simply a point query using an index, similar to other queries in an OLTP system.

The only limitation I see is scalability. It's not easy to make PostgreSQL distributed, but solutions like Citus exist, making it still possible.

(I'm the author of pgvecto.rs)

VoVAllen··on How We Made PostgreSQL a Better Vector Database
Hi, we've solved the problem you mentioned! Please take a look on our open source postgres vector extension https://github.com/tensorchord/pgvecto.rs.

Our index building process is significantly faster than pgvector on hnsw because we can utilize all the cores, whereas pgvector can only use one core. And for the filter support, we do support pre-filtering, which will guarantee enough results no matter the condition is.

VoVAllen··on Show HN: Epsilla – Open-source vector database with low query latency
Why it's a yellow flag?
VoVAllen··on Show HN: Epsilla – Open-source vector database with low query latency
Why did you choose SpeedANN instead of other new indexes such as DiskANN? And you changed the color of epsilla in every benchmark figure, which is quite confusing
VoVAllen··on Show HN: 20x faster pgvector alternative written in Rust
- We used a different index method called HNSW, which is more widely used in vector search area and also faster than ivfflat used by pgvector.

- There are some drawbacks with HNSW. It's designed for memory usage but not so suitable as a disk database index. Currently the storage of the HNSW index is managed outside postgres's buffer system, which is not so ideal. We're exploring more indexing method such as DiskANN to see whether we can integrate with postgres more. Fully integration with postgres can make things work like a charm

VoVAllen··on Launch HN: Resend (YC W23) – Email API for developers using React
Your homepage looks quite similar to modal.com
VoVAllen··on Why do tree-based models still outperform deep learning on tabular data?
It's all about features and data scale. Recommendation system itself is actually a large table, DL method already proved effectiveness there. Let's say if you have text in your tabular data. Tree model(with traditional method such as tfidf) will do much worse than the transformer-based model. DL always suffers from inadequate data, so if there's no enough data or inductive biases, tree model can be a better choice in that way.
← PreviousPage 2 of 2