Pinecone raises $100M Series B
pinecone.io
pinecone.io
The key thing is that it's in-memory and allows you to combine attribute-based filtering, together with nearest-neighbor search.
We're also working on a way to automatically generate embeddings from within Typesense using any ML models of your choice.
So Algolia + Pinecone + Open Source + Self-Hostable with a cloud hosted option = Typesense
Supabase was also asking for sparse vectors https://github.com/pgvector/pgvector/issues/81
Speaking of the repo, they have a number of features they want to add if anyone is interested in contributing, there's lots of room for advancement. Many of these features already have active branches https://github.com/pgvector/pgvector/issues/27
A marqo.ai dev is currently working on adding HNSW-IVF and HNSW support https://news.ycombinator.com/item?id=35551684 and the maintainer has recently noted that they are actively working on an IVFPQ/ScaNN implementation https://github.com/pgvector/pgvector/issues/93
The pgAnn creator actually asked about performance a month ago here https://github.com/pgvector/pgvector/issues/58
Expect to see performance improve dramatically later this year.
recent vector database fundraises:
- Chroma - $18M seed https://www.trychroma.com/blog/seed
- Weaviate - $50m A https://www.theinformation.com/articles/index-ventures-leads...
- Pinecone - $100M B
Harder to set up than wrappers like Chroma, but very powerful.
Maybe I'm not the target audience but after spending some time poking around I couldn't honestly couldn't even figure out how to use it.
Even a simple Goole Search for "What is a Vector Database" ends up with this page.
https://www.pinecone.io/learn/vector-database/#what-is-a-vec...
Pinecone is a vector database that makes it easy for developers to add vector-search features to their applications
Um okay... what's "vector-search"? For that sake what the eff is a "Vector" to begin with? Finally after getting about a third down the page we start defining what a vector is....
Maybe I'm not their target audience but I ended up poking around for about an hour or two before just throwing up my hands and thinking it wasn't for me.
Ended up just sticking with Algolia, since we had them in place for Search anyway...
When they say “vector-search” they mean semantic search. I.e. “which document is the most semantically similar to the query text”.
So how do we establish semantic similarity?
In a database like Elasticsearch, you store text and the DB indexes the text so you can search.
In a vector DB you don’t just store the raw text, you store a vectorized version of the text.
A vector can be thought of as an array of numbers. To get a vector representation we need some way to take a string and map it to an array while also capturing the notion of semantics.
This is hard, but machine learning models save the day! The first popular model used for this was called “word2vec” while a more modern model is BERT.
These take an input like “fish” and output a vector like [ 3.12 … 4.092 ] (with over a thousand more elements).
So let’s say we have a sentence that we vectorized and we want to compare some user input to see how similar it is to that sentence. How?
If we call our sentence A and the input vector B, we can compute a number between zero and one that tells us how similar they are.
This is called cosine similarity and is computed by taking the dot product of the two vectors and dividing by both of their magnitudes.
When you load a bunch of vectors in a vector DB, the principal operation you will perform is “give me the top K documents that are similar to the input”. The databases indexing process computes k nearest neighbors algorithm on all vectors in the DB and stores this for use at query time.
Without the indexing process there is no real difference between a vector db and key value store.
I wasn't looking for one ;-) I was looking for a recommendation engine, similarly most often I'm looking for various ways to use ML and AI to improve various features and workflows.
Which I guess is my point, I don't know who Pinecone's target market is but from following this thread it seems like all the folks who know how to do what they do have alternatives that suit them better. If they are targeting folks like me they're not doing it well.
Pinecone's examples[1] (hat tip to Jurrasic in this thread - I've seen these) all show potential use cases that I might want to leverage, but when you dive into them (for example the Movie Recommender[2] - my use case) I end up with this:
The user_model and movie_model are trained using Tensorflow Keras. The user_model transforms a given user_id into a 32-dimensional embedding in the same vector space as the movies, representing the user’s movie preference. The movie recommendations are then fetched based on proximity to the user’s location in the multi-dimensional space.
It took me another 5 minutes of googling stuff to parse that sentence. And while I could easily get the examples to run I was still running back and forth to Google to figure out what it was doing in the examples - again the documentation is poor here. I'm not a Python dev but I could follow it but I still had to google tqdm to figure out it was a progress bar library?
Also, and this is not unique to Pinecone, I've found generally that while some things are fairly well documented on "Here's how to build a Movie Recommender based on these datasets) frequently in this space there's very little data on how to build a model using your own datasets ie how to take this example and do it with your own data.
emb = model(text)
Now you got the embedding. What can you do with it? you can calculate how similar it is to other texts. emb1 = model(text1)
emb2 = model(text2)
similarity = sum([a * b for a, b in zip(emb1, emb2)])
Just a multiply and add, this is trivial! So if you do that for a million texts, you got a search engine. Vector DBs are automating this for you. There are free libraries just as good. And free models to embed text with, OpenAI also have some great embeddings. You can use np.dot to compute similarities fast, up to 100,000 vectors it's the best way and get exact, not approximate results.The great thing about embedding text is the simplicity of the API and the similarity operation. It's dead simple to use. You can do clustering, classification, near neighbour search / ranking, recommendation, or any kind of semantic operations between two texts that can be described as a score. If you cache your vectors you can search very very quickly with np.dot or other methods, in a few ms. Today you can also embed images to the same vector space and do image classification by taking the text label with max dot product.
You can also train a very small model on top of embeddings to classify the input into your desired classes, if you can collect a dataset. Embeddings are the best features for text classification. You can think of this embedding method as a way to slice and dice in the semantic space like you do with strings in character space. All fast and local, without GPUs.
I'm shocked.
Faiss is a collection of algorithms for in-memory exact and approximate high-dimensional (e.g., > ~30 dimensional) dense vector k-nearest neighbor, it doesn't add or really consider persistence (beyond full index serialization to an in memory or on disk binary blob), fault tolerance, replication, domain-specific autotuning and the like. The "vector database" companies like Pinecone, Weviate, Zilliz and what not will add these other features to turn them into a complete service, they're not really the same. pgvector seems to be DB-backed IndexFlat and IndexIVFFlat (?) from the Faiss library at present but is of course not a complete service.
However which kind of approximate indexing you want to use very much depends upon the data you're indexing, and where in the tradeoff space between latency, throughput, encoding accuracy, NN recall and memory/disk consumption you want to be (these are the fundamental tradeoffs in the vector search domain), and whether you are performing batched queries or not. To access the full range of tradeoffs you'd need to use all of the options which are available in Faiss or similar low-level libraries which may be difficult to use or require knowledge of underlying algorithms.
(I'm the author of the GPU half of Faiss)
Is it possible they are confusing the use of embeddings across whole swaths of text to do a semantic search with the embeddings that happen on a per token basis as data runs through an LLM? Same word, same basic idea, but used so differently that they may as well be different words?
The alternative - and I believe the way the ChatGPT web app currently works - is just to stuff/mapreduce user input and machine response into the context window on each step, which quickly gets quite lossy.
I consider vector engines to be "hot" models, given they are storing the vector representations of text already run through the "frozen" model.
Having written something a while back that indexes documents and enters into discussion with them, I'm pretty sure ChatGPT is using some type of embedding lookup/match/distance on the history in the window. That means not all text is submitted at the next entry, but whatever mostly matches what is entered by the user (in vector space) is likely pulled in and sent over in the final prompt.
Using semantic search to find relevant chunks seems misguided but practical in the short term. One of the key benefits of LLMs is they can take into account a lot of context.
But summarization is better to keep the model on topic for most cases. And there are other tricks.
Vectors and semantic search are one (likely questionable way given LLMs can likely reason over a table of contents or similar better) to search a large corpus or very large document. It's really only appropriate for a specific set of use cases. It's not some "general memory layer" for AI.
Disclaimer: I work for a16z and on the infra team, so consider me biassed.
As for a corpus of documents (which is what you are presumably talking about), there are a couple problems with what you are saying:
First, you are implying that the content is always new - that's not true for many cases folks are talking about solving (like technical support or customer support), so it's a one time fee to summarize the corpus. You might run it periodically for updates.
Second, there is an assumption that a basic semantic search is the best way to search documents to find the most relevant content. That's questionable before the existence of LLMs, but with LLMs you are basically assuming your cosine similarity search on your vectors is better than an LLM can do with a simple table of contents and question "where should I search?" I haven't seen someone do a detailed study, but the implicit assumption that semantic search is the best idea for text could easily be a bad one.
Third, it assumes the quantum of data to search through is astronomically large and/or getting bigger compared to almost certain decreases in inference cost and increases in input tokens. This will be true for some subset of things, but unlikely to be many and in the cases it is true they'll do something more sophisticated than embeddings and embedding search. They'll probably fine tune the underlying model on an ongoing basis.
Regardless - the post you guys wrote seems... like a stretch for a definition of what this really is And, at least on the surface vector databases appear to be commodity infra. Pinecone might be growing fast now, but how do they ever make much money above their costs? But, you guys seem smart, so maybe there is something there?
I don't quite understand how general summarization would work. If you use an LLM to simply to summarize in order to feed it into a prompt, the summarization needs to be specific to the query. i.e. "summarize what this text says about topic X". You can't summarize long text in a generic way without losing information. Or do I misunderstand the comment?
If you have a perfect table of context (or better, an index by topic) you may not need semantic search. But for the typical use case we are seeing you have unstructured data without an index (e.g. tech support knowledge db entries, company reports, emails). For that, semantic search work quite well.
For the sizes, the observation is that the data that people want to search over (e.g. your email, a wiki, JIRA, a knowledge base) is far larger than the context length. You are correct that we assume that inference cost and speed won't decrease sufficiently quickly in the near future. Why is a longer topic, but in a nutshell GPU speed increase is ~2.5x gen/gen and other than overtraining vs. Chinchilla we don't see immediate model gains. But that is speculative, we don't know what's in store.
To some degree we are just reacting to user adoption in the market. We don't build these systems, but if we see enough of them eventually we recognize the pattern. And while I am optimistic, we could be wrong. AI is major revolution and we are all students.
edit: disclaimer, I work for a16z.
edit: To me this is a better summary of what a vector db is useful for: https://cloud.google.com/blog/topics/developers-practitioner...
And if someone is building a chat interface which is effectively a search product then they are going to find these things useful. But it's not a generic LLM memory layer or something.
Admittedly I’m hand-waving a bit around “relevant” and “irrelevant” - clearly your vector search setup has to be fit for purpose. That’s a talent all on its own, so I wonder if we will see competing approaches at the vectorstore level or if it’s relatively settled. Anyway, I’m out of my depth at that point so I’ll leave it there.
I agree that a combined approach is likely useful.
https://python.langchain.com/en/latest/modules/memory/types/...
It also has a more basic version that just keeps a log of past messages.
I don't know whether there's a way (or even a need) to combine these approaches. In a long conversation, it might be useful to trust more recent information more than earlier messages, but Langchain's vector memory doesn't care about sequence.
It quite angers me that people (on HN) will consider the following to be benefits worth mentioning as pros to the consumer:
>Sleek/shiny finish
>Marketing/Branding
>Ability to Monetize
We arent shareholders, all 3 of these are bad for the customer.
Every new software company like this has "why wouldn't everyone just use x existing open source project, why even try to make it a real business with a hundred devs, actual support/marketing, and big ambitions to be more than a plugin to Postgres?"
Based on the videos and interviews with their lead dev I've seen Pinecone has some quite large plans by integrating with a wider stack and integrating with company databases, well beyond what they have done so far releasing an early version of the DB.
Regardless, getting wider adoption via actual businesses investing in marketing/sales to seed ideas in the market can spur development and potentially progress/innovate the tooling across the wider market, that feeds back into open source.
From what I've seen, the big limitation currently is dimensionality. Most of the more advanced models have a high dimensionality and especially Elasticsearch and Lucene limit the dimensionality to 1024. E.g. several of the openai models have a much higher dimensionality. Opensearch works around this by supporting alternate implentations to lucene for vectors.
Of course it's a sane limitation from a cost and computation point of view, having these huge embeddings doesn't scale that well. But it does limit the quality of the results unless you can train your own models and tailor them to your use case.
If you are curious on how to use this stuff, I invested some time a few weeks ago getting my kt-search kotlin library to support this and wrote some documentation for this: https://jillesvangurp.github.io/kt-search/manual/KnnSearch.h.... The quality was underwhelming IMHO but that might be my complete lack of experience with this stuff.
I have no experience with pinecone and I'm sure it's great. But I do share the sentiment that they might not come out on top for this. There are too many players here and it's a fast moving field. OpenAI just majorly moved the whole field forward enormously in terms of what is possible and feasible.
It "just works".
You can always use an open-source alternative. But you probably shouldn't self-host at early stage atleast.
https://towardsdatascience.com/milvus-pinecone-vespa-weaviat...
Isn't there still a token limit as to how much ChatGPT can hold in working memory?
Is the goal that ChatGPT can query the vector database directly to get information out, and if so, how is that different than using a regular database?
So the sentence "I started working as a programmer" will be very close to "I began my job as a software developer". This makes it very powerful for natural language search.
So when the user asks a bot "Find the text message John sent me 3 years ago about wanting to found a company, I think it was like a hang gliding company? or parasailing? idk"
Behind the scenes you can ask GPT "Output a list of 10 candidate sentences that are plausible text messages that John may have sent", and it spits out
"I'm thinking of starting a hang gliding company" and "I might found a parasailing company" etc.
Then you query the vector DB for those imaginary sentences, and as a result you get sentences that are semantically similar. You take the top N nearest neighbors and plug them in to the GPT context window and say:
"Here are 100 sentences that might match the original query. If any of them match, select the number that matches. Otherwise output null"
If you engineer this system well, you can get pretty decent results.
This is great for key:value type querying, but IMO there is a lot of ground to be explored by extending it with more graph-like links. I.e. use vector search to get some initial nodes that are better than random, but then start doing a little beam-search algorithm from those nodes to find nodes they are "linked" to in some way that may answer the query better.
Does it work for text data or can it work for other types of data as well?
It all depends on how you produce the vectors before storage. The vector database just stores them.
Only a sucker being forced to by their investors would use pinecone.
Who is even paying?
The homepage offers a few clues: Shopify, Gong, Zapier, HubSpot, Expel, and several thousand others. That includes huge enterprises who tend not to want their names shown publicly.
Basically there are many companies with tens of millions, hundreds of millions, and even billions of embeddings. If they care about performance and reliability, and don't want to tie up an entire team of engineers to manage a self-hosted solution, then Pinecone makes a lot of sense for them.
In a way this also answers the many questions about "Pinecone vs [whatever]" ... If you're dealing with <1M embeddings the differences between your options will hardly matter — just pick whatever's easiest for you. If you're already using a managed DB that introduced something that's good enough for you... just use that. Though we still work hard to make Pinecone the easiest choice and have features that many basic solutions don't have, such as hybrid search (sparse + dense vector embeddings) for better search results.
Can someone explain what the use case is for vector DBs like pinecone, milvus etc. vs a fully featured search engine like Vespa, ElasticSearch etc. which also support vector search features?
Is there something about running this type of index operationally that is particularly difficult?
Overall I’m really enjoying the tech.
I closed the browser tab here. It's a usual techbro hustle, nothing to see here, move along. Hype, whatnot, couldn't care less. I am sure they will make money but it'll be detrimental to society. As always.