Vector Search with OpenAI Embeddings: Lucene Is All You Need
arxiv.org
arxiv.org
pgvector is even supported out of the box on Azure and AWS RDS.
Just spin up a docker container [1], add a vector column to your table and you're ready for embedding search.
[1] https://hub.docker.com/r/ankane/pgvector
If you're starting out with a prototype - do yourself a favour, steer clear of the chromadb examples with langchain. In fact steer clear of langchain in general :) Just go for OpenAI API and PostgreSQL+PGVector - you'll have to do some boilerplate - but the stuff in langchain is just terrible, you'll have to rewrite it and do the boilerplate at some point anyway and this stack is super simple to deploy.
Even with pgvector, there's no good way to write a simple tutorial for embedding newbies.
I'm coming at this from a perspective of a competent dev trying to build a tool with OpenAI APIs - which, from what I can tell, is a growing topic.
If you're experienced with building APIs and new to the LLM stuff - skip the langchain and chromadb nonsense. Just use the OpenAI APIs and pgvector.
Chromadb and langchain are usefull when writing notebook prototypes to get an idea of how this stuff works - but discard immediately after that phase and save yourself the trouble of porting later.
MySQL is very attached to /etc/mysql, /var/log/mysql and /var/lib/mysql, which makes sense if one thinks of it as a piece of a distribution and makes no sense if one thinks of it as a service that stores data in a filesystem or directory that one sets up for the purpose. Apparmor and (don’t get me started) SELinux dig this in deeper. What if you want two MySQLs on one host? What if you don’t want to mount something on /var/lib/mysql?
If mysqld were invoked by pointing it at a configuration and data and it just worked, I’d be more okay with it.
The fact that I really don’t want to couple upgrades of MySQL to distro upgrades is just icing on the cake.
For development, perhaps. For production, absolutely not.
I wish HN had some bot that would just delete any comment from people recommending installing databases. Because 99.9% of the time it's from those who have no experience in running one in a Production environment. Keeping it secure and ensuring backup/restore cycle works is seriously non-trivial.
Would love to know how running a database is considerably less work than adding a library to your existing app.
Managed Postgres is a thing, probably much more common than managed lucerne. I can name three managed Postgres providers off the top of my head (AWS, Vultr, Supabase).
There are search engines e.g. Solr, ElasticSearch but those aren't what this paper is proposing.
To be clear I object to this:
> I wish HN had some bot that would just delete any comment from people recommending installing databases.
Most applications already use a database. Pooh-poohing Postgres as a solution is silly. For many people, it’s actually going to be less complex to take to production than installing a library, especially one with its own complex storage needs.
Also: technically, isn’t Lucerne a database?
I think if you're past pgvector performance you won't be listening to a random guy talking about pgvector but have a good understanding of the space.
If you're new (like I was a few months ago) save yourself the time I wasted on the noobtraps I mentioned. It scales way better than the OpenAI API for my use cases.
Lucene (or rather elastic/open search) is way overkill for my needs
AFAICT standalone optimized vector stores will still have their place, like super latency sensitive or high-throughput scenarios. Unfortunately for many stakeholders, venture-scale returns to justify the megarounds from the peak vc fomo of the last few years seems unclear. The TBD hail mary here may be generative AI: even if most modern vector search workloads are largely fine with regular DB extensions to support a few kinds of vector indexes, the continued growth of generative AI, knowledge graphs, etc., may somehow grow the pie for non-standard DBs here. That's not obvious to me. For example, with Databricks LakehouseIQ, a lot of the use case may be eaten by the data warehouses.
I am currently using it this way and it is really easy to just get started.
I don't know how well it compares with others, however.
IMHO, it's a gigantic self-own and doesn’t promote Lucene in a good way. For example, by demonstrating how they get only 10 QPS out of a system with 1TB of memory and 96 v-cpu's (after 4 warmups).
The HNSW implementation in Lucene is fair, and within the same order of magnitude as others. But, to get comparable performance, you must merge all immutable segments to a single segment, which all Lucene oriented benchmark does, but which is not that realistic for many production workloads where docs are updated/added in near real-time.
it really depends on how real time you need the search to be tho.
What i've seen is a green/blue lucene index. The updates happen on one (let's say the blue), while searches happen on the other (green). The segment merging happens periodically for the blue (or even smarter, let's say, after some known amount of time and updates combined), and then the index are switched. Depending on how often new documents come in, and "real time" you need, this may be sufficient.
Attention is all you need was a breakthrough paper. It fundamentally changed the ML landscape and got us out of a huge roadblock with rnns.
If you seriously think you have something similarly impactful on your hands, then sure go ahead with that name. But there’s been a bunch of papers where I found it distasteful. At best its just not funny. But this isn’t even really much of a paper. I’ve seen blog posts with more substance. Hell, even YouTube videos.
I don’t know, I guess I just don’t really get the joke.
I don't think referencing well known papers has ever (well maybe not, ever) implied that the authors feel their work is on par with original. It's a pretty common practice in some academic circles when there is an impactful paper with a catchy name and you simply want to pay a bit homage and have a less boring title than you might have otherwise.
txtai (https://github.com/neuml/txtai) can build indexes with Faiss, Hnswlib and Annoy. All 3 libraries have been around at least 4 years and are mature. txtai also supports storing metadata in SQLite, DuckDB and the next release will support any JSON-capable database supported by SQLAlchemy (Postgres, MariaDB/MySQL, etc).
"We provide a reproducible, end-to-end demonstration of vector search with OpenAI embeddings using Lucene on the popular MS MARCO passage ranking test collection......This suggests that, from a simple cost-benefit analysis, there does not appear to be a compelling reason to introduce a dedicated vector store into a modern "AI stack" for search, since such applications have already received substantial investments in existing, widely deployed infrastructure."
Curious why stop there, why even use OpenAI embeddings and not use, say, LLaMA embeddings and create a truly open stack.
The embeddings are just the "data" that's in the database. Swapping out getting embeddings from OpenAI with Llama is as trivial as putting information about your own customers in your own database as opposed to using info on someone else's customers.
Also, all embeddings are basically equivalent for this use case.
I also do not understand why VCs are investing in this space. The base case is that they are almost completely interchangeable c/o langchain and other intermediaries.
I do understand that some vector databases have strengths over others w/r/t scaling OUT, however, they do not have stickiness and time erodes the scaling advantages via both competitors catching up and via machines getting cheaper allow for scaling UP
The history of commercial DBs was usually supported by a variety of use cases, proprietary hooks to keep customers, choices on the CAP theorem, etc. Almost none of that applies here given the minimal interaction modes we have with vector DBs.
Could anyone speak to the case to invest in vector DBs?
(and to address the parent take, "elastic/redis/opensearch/vespa or with postgres/mongodb/oracle/mysql" - this is one of the most crowded spaces in the marketplace and i have no idea why customers would choose an upstart for the sake of consolidation rather than clear winners in a best of breed solution.
Late last century Oracle felt urgency to compete with the new hotness back then, object databases. The guys in charge of the Oracle database itself pushed back, not too much was done, and in the end those competitors all flamed and died.
Object databases gave way to graph databases. They broke through a little but RDBMS continued to rule.
Then came the NoSQL movement. RDBMS vendors ended up adding json columns and the pure NoSQL vendors flamed and died (excepting Mongo DB which is web scale (https://www.youtube.com/watch?v=b2F-DItXtZs)).
Thus it will continue forever.
MS MARCO is a well known benchmark that mirrors a pretty common use case: https://www.sbert.net/docs/pretrained-models/msmarco-v3.html
https://huggingface.co/spaces/mteb/leaderboard
To intentionally oversimplify, embeddings fall into 2 main categories if you do a cursory search:
- Good at semantic similarity: Most common case, check if two strings have similar meaning, even if the words don't match exactly.
- Good at Q/A: Finds text that can answer a question. Sound similar to semantic, but "What is a dog" and "What is a cat" are very similar sentences but very different questions. These models cluster questions closer to their answers and further from other questions. They also handle the difference in length between questions and answers better.
(Some leading embedding models use prefixes during training to let you adjust performance between those two tasks on the fly)
LLaMA embeddings weren't optimized to any specific task, you'd just be hoping that they tangentially align with some arbitrary goal.
OpenAI had to fine tune their embeddings model, and despite being massively oversized you can see it's not at the top of the leaderboard compared to much smaller models.
It's also available opensource: https://github.com/marqo-ai/marqo
A voice in my head seems adamant that the solution to this whole space of problems is neatly managed by one clever schema and minimal computational resources. It has only been growing louder and more confident in this over time.
Finer points:
1. I remember FAISS is very accelerated for GPUs. How does Lucene compare there?
2. "we’re not convinced that enterprises will make the (single, large) leap from an existing solution to a fully managed service" --> Fair point but not everyone uses Lucene? I feel it is weird that this "existing solution" (Lucene) is assumed to be already adopted.
We'll be demo'ing an embedding service that uses Instructor Large/XL embeddings + GPT-4 keyterm extraction this next week.
So you don't have another component that you need to cost, integrate, provision, manage, secure, backup etc.
"We had to incorporate logic for error handling in our code, given the high-volume nature of our API calls"
This just seems like an asinine thing to add to a technical paper. "We had to handle errors..."
The tricky part is that 10% of your code will become the basis of your thesis / postdoc;)
… actually through codecs Lucene has a whatever-you-want-to-build implementation
https://github.com/search?q=repo%3Aapache%2Flucene%20pinecon...