Pinecone integrates AI inferencing with vector database
blocksandfiles.com
blocksandfiles.com
After reading the article, it seems Pinecone just now supports in-DB vectorization, a feature that is shared by:
- DataStax Astra DB: https://www.datastax.com/blog/simplifying-vector-embedding-g... (since May 2024)
- Weaviate: https://weaviate.io/blog/introducing-weaviate-embeddings (as of yesterday)
Weaviate seems to have added a similar capability — kind of wild that they announced on the same day.
Looks like Pinecone also includes reranking as part of the same process — did Weaviate add that as well?
> Astra DB seems to just be a tutorial showing how to generate embeddings using another service.
The link I shared showed how a single request to Astra DB's data API has Astra DB automatically create embeddings behind the scenes, integrating with an embedding service the user chooses when they set their database up. Indeed embeddings are generated by another service and not in-house, but from an end-user perspective, they don't need to generate embeddings themeselves as was the prior art and coordinate requests between:
- get text - generate embeddings - take embeddings and send to DB
As of May when they announced Vectorize, one request did all that. I believe from an end-user experience, this is really analogous to what Weaviate and Pinecone are offering unless I'm missing something.
Timescale most recently added it but, yes a bunch of others: Weaviate, Spice AI, Marqo, etc.
https://qdrant.tech/documentation/concepts/hybrid-queries/
And handles embedding creation with its fastembed package.
Makes a lot of sense to me to combine embedding, retrieval and reranking — I can imagine this being a way that they can differentiate themselves from the popular databases that have added support for vector search
I assumed that a specific flavour of LLM was needed, an “embedding model” to generate the vectors. Is this announcement that pinecone is adding their own?
Is it better or worse than the models here: https://ollama.com/search?c=embedding For example?
> Is this announcement that pinecone is adding their own?
TLDR: they trained their own embeddings model and rely on Cohere for ranking. Pinecone (the database) uses this model automatically to generate and store embeddings.
> I assumed that a specific flavour of LLM was needed, an “embedding model” to generate the vectors.
You're mostly right, with one caveat: embeddings models aren't really LLMs in that they're not very large: they just map semantic meaning to numerical space.
> Is it better or worse than the models here: https://ollama.com/search?c=embedding For example?
This is the golden question. As far as I know, there is no appropriate benchmarking/eval data about this. I think the real value is the first-class integration between their model and their service.
This is building it into the vector DB such that you send it the content and it is "built in".
Seems silly. It's like bundling a stove with cookware. But cookware fit specific niches and have different life cycles. I get that it might cater to some "drop in solution" targets, but seems of no value for most engineered, long-term solutions.
I've played around with Weaviate & Astra DB but Marqo is the best and easiest solution imo.