But I want to comment on another thing I often hear: "You don't need a vector database - just use Postgres or Numpy, etc". As someone who moved to Pinecone from a Numpy-based solution, I have to disagree.
Using a hosted vector database is straightforward. Get an API key from Pinecone, send them your vectors, and then query it with new vectors. It's fast, supports metadata filtering, and scales horizontally.
On the other hand, setting up pgvector is a hassle - especially since none of the Cloud vendors support it natively, and a Numpy-based solution, while great for a POC, quickly becomes a hassle when trying to append to it and scale it horizontally.
If you need a vector database, use a vector database. You won't regret it.