Pgvector: Fewer Dimensions Are Better
supabase.com
supabase.com
There are some trade-offs (it was trained on English-only text), but this is a good indicator of the direction where embeddings might go (code-specific, language-specific, etc)
[0] https://huggingface.co/Supabase/gte-small
[1] MTEB leaderboard: https://huggingface.co/spaces/mteb/leaderboard
I would love to see a comparison for some typical use cases using various methods of chunking input documents.
For example in chatgpt-retrieval-plugin[0] repo default chunk size is just 200 tokens
this is anyway a limitation, no doubt, but chunking is pretty often used
[0] https://github.com/openai/chatgpt-retrieval-plugin/blob/main...
fwiw, 98% of Supabase project use OpenAI for embeddings, so a lot of this content is just to help devs identify the best options they have available
And the model is only 70MB!
btw, you can find the dataset with embeddings generated by all 3 mentioned: text-embedding-ada-002, all-MiniLM-L6-v2, and GTE-small on huggingface[0]
and big thanks to Stephan Sturges for his dataset[1]. we just extended his OpenAI ones and texts themselves with oss ones
[0] https://huggingface.co/datasets/Supabase/wikipedia-en-embedd...
[1] https://www.kaggle.com/datasets/stephanst/wikipedia-simple-o...
Our focus with Zep is on LLM app use cases, and in particular, searching over chat histories and documents for RAG apps. You can, however, use Zep to turn any Postgres instance into a vector store with great developer experience. pgvector index configuration and query tuning can be challenging. We've tried to do much of this work for developers.
Also, I'm super excited about the prospect of HNSW support in pgvector, slated for pgvector 0.5.[2]
[0] https://github.com/getzep/zep
[1] https://blog.getzep.com/introducing-the-zep-document-vector-...
[2] https://github.com/pgvector/pgvector/issues/181#issuecomment...
The highest ranking is `gte-large`, which has 1024 dimensions - larger than `gte-small`(384), but still smaller than `text-embedding-ada-002` (1536)