Find anything fast with Google's vector search technology
cloud.google.com
cloud.google.com
Running vector search (also sometimes referred to as semantic search, or a part of semantic search stack) is a trivial matter with open-source libraries like Faiss https://github.com/facebookresearch/faiss
It takes 5 minutes to set up. You can search billion vectors on common hardware. For low-latency (up to couple of hundred milliseconds) use cases, it is highly unlikely that any cloud solution like this would be a better choice than something deployed on premise because of the network overhead.
(worth noting is that there are about two dozen vector search libraries, all benchmarked at http://ann-benchmarks.com/ and most of them open-source)
A much more interesting (and harder) problem is creating good vectors to begin with. This refers to the process of converting a text or an image to a multidimensional vector, usually done by a machine learning model such as BERT (for text) or ImageNet (for images).
Try entering a query like 'gpt3' or '2019' into the news search demo linked in the Google's PR:
https://matchit.magellanic-clouds.com/
The results are nonsensical. Not because the vector search didn't do its job well, but because generated vectors were suboptimal to begin with. Having good vectors is 99% of the semantic search problem.
A nice demo of what semantic search can do is Google's Talk to Books https://books.google.com/talktobooks/
This area of research s fascinating. For those who want to play with this more, an interesting end-to-end (including both vector generation and search) open-source solution is Haystack https://github.com/deepset-ai/haystack
What was the largest index you've had on Faiss? That seems to affect whether people think of it as more than adequate or terribly inadequate.
I compared BERT[1], distilbert[2], mpnet[3] and minilm[4] in the past. But the results I got "out of the box" for semantic search were not better than using fastText, which is orders of magnitude faster. BERT and distilbert are 400x slower than fastText, minilm 300x, and mpnet 700x. At least if you are using a CPU-only machine. USE, xlmroberta and elmo were even worse (5,000 - 18,000x slower).
I also love how fast and easy it is to train your own fastText model.
[1]: distiluse-base-multilingual-cased-v2
[2]: multi-qa-MiniLM-L6-cos-v1
[3]: multi-qa-mpnet-base-cos-v1
[4]: multi-qa-MiniLM-L6-cos-v1
A concrete example is DPR which is a state of the art dense retriever model for wikipedia for question answering, when applying that model on MS Marco passage ranking it performs worse than plain BM25.
If you do have a tiny index and want to try Google's version of vector search (as an alternative to Faiss), you can easily run ScaNN locally [1] (linked in the article, that's the underlying tech). On small scale I had better perf with ScaNN
[1] https://github.com/google-research/google-research/tree/mast...
But I get 1/5 - 1/10 hit ratio(successful/empty searches). That's not habit forming, memory forming for me.
Is there a core use case where I would get a good hit ratio ?
A similar concept was shown here recently: https://search.marginalia.nu
Haystack does not really compare to huggingface as they are directly using it as a library, it's more of a layer on top.
Serving a search system for 100K lines (short document?) is trivial on a single CPU machine, so the costs will be low.
For training if you stick to fine tuning and your dataset is relatively small (<10M) you could use collab with a free gpu. Training time really depends on your use case, the simplest classification may be handled in a couple of hours.
(1) Google's offering returns with-in <5ms, in my experience. (2) the demo is for paragraphs, not short text. You're putting mismatched data into the input, of course it's not going to work. Try a paragraph as suggested.
The demo featured in this PR takes about 800-1000ms total to produce search results. How much of that is the actual API is not known. Typically, https request to an API in the cloud will cost you at least 50ms of network latency, more likely 100ms-200ms. If you are running vector search on premise you will obviously not have this overhead.
Text embeddings typically work for short text as well as paragraphs (paragraph embeddings are usually mean/max of word embeddings anyway) simply because most commercial use cases demand handling of short text input (because nobody is inputting a paragraph into a typical search box; what use is a news search if you can not type a single word like 'Biden' or 'gpt3' into it).
The offering is similarity search, not a search engine. They offer image to image as another comparison point.
I was obviously talking about a general use case where a user considers using an API like this vs running Faiss, and their server can be anywhere (a use case that is more common to me personally).
Typical Google Cloud Bigtable P(50) is <= 4ms to a GCE VM in the same region.
Source: I work on Google Cloud Bigtable.
EDIT: I should clarify that applies to point reads on an established connection using gRPC-over-HTTP/2 from e.g. Java/golang/C++ client, and doesn’t include cold connection set-up / uncached auth / TLS negotiation, but those are well into P(99.999) territory if you’re doing millions of QPS (which applies to most of our large customers).
1. Incrementally updating the search space. Not that easy to do, and becomes more important to not just do the dumb thing of retraining the entire index on every update for larger datasets.
2. Combining vector search and some database-like search in an efficient manner. I don't know if this Google post really solves that problem or if they just do the vector lookup followed by a parallelized linear scan, but this is still an open research/unsolved problem.
And you’re right, it wasn’t easy to build.
On your second point about efficient filtering, check out this article I wrote outlining how filtered vector search works in Weaviate: https://towardsdatascience.com/effects-of-filtered-hnsw-sear...
For even more details on filtering, check the documentation: https://www.semi.technology/developers/weaviate/current/arch...
And even if you do just want ANN and nothing else, some people just want to make API calls to a live service and not worry about anything else.
Anyway, for me, the number one priority was latency and it is hard to beat on-premise search for that.
Even then, a vector search API is just one component you will need in your stack. You need to pick the right model, create vectors (GPU intensive), then possibly combine search results with keyword based search (say BM25) to improve accuracy etc. I am still waiting to see an end-to-end API doing all this.
It's an interesting point tho. Maybe it's good to add this to the docs as well
That’s kinda the idea of Weaviate. You might like the Wikipedia demo dataset that contains all this. You indeed need to run this demo on your own infra but the whole setup (from vector DB to ML models) is containerized https://github.com/semi-technologies/semantic-search-through...
Vespa.ai supports combining dense vector search with keyword search and ranking, see https://docs.google.com/presentation/d/1vWKhSvFH-4MFcs4aNa9C...
There is also a Vespa sample application (open source, Apache 2) demonstrating multiple different retrieval and ranking strategies over at https://github.com/vespa-engine/sample-apps/blob/master/msma...
You should copy paste a part of an existing article. It will embed it into a multidimensional space and do a similarity search (the same way it was trained to (by converting a full article paragraph to a vector).
If you give it just a word it can't convert it to a meaningful representation because the network wasn't trained to do this.
But you can train it differently and have it able to handle a few words. For example you can summarize every article to a few sentences and keywords and use a traditional keyword search.
One usual way to create good vector representation is to use encode simultaneously two different space to the same vector space. You encode 'queries' and 'answers' such that they are close for the known (query-answer) pairs. This is what CLIP did, encoding both images and their corresponding description to a same vector space.
You can download the precomputed clip embeddings LAION-400-MILLION OPEN DATASET on academictorrents.com
CLIP can do such thing for the problem of semantic image search because the problem of matching an image to its description is quite well defined. But quite often there is no unique apriori meaningful way of matching a query to an answer, specially as the index get big.
In the case of a basic query 'gpt-3', the query is quite vague and its not obvious with respect to which direction you should do the ranking (Do you meaning you want articles generated by gpt-3 ? Articles containing gpt-3, a basic definition ?). There is no a priori good answer, and that's where you can use your additional context to refine the query. For example Siri or a NLP bot, could ask you to be more explicit in what you mean.
Or it can have multiple representation of your vector space and return the top-1 for each of those representation, and hope that you give it feedback by clicking the one that was more meaningful to you as requester.
Disclose: I have built a vector search engine to proof this idea[2]
Indeed, this is the hardest problem. Vector search shines when used in-domain using deep representation learning, for example bi-encoders on top of transformer models for text domain.
However, these models does not generalize well when used out of domain as the representations changes. Hence, in many cases, simple BM25 beats most of the dense vector models when used in a different domain. See https://arxiv.org/abs/2104.08663
Do sentence, paragraph, or full doc vectors work the best? Do things work better creating vectors from sentence vectors for longer sections, or have you had better results making longer vectors directly? Do different vectorizing algorithms work better on different sized sections of text?
I have not had much luck finding discussion or write ups on this type of stragetizing on chunks and algorithms to apply, beyond just people saying that yes that's how it's done.
What I was attempting to ask was more on how to effectively compare longer sections of text. It sounds like you are saying it is fine to compare them directly. So, due to very, very irregular formatting between texts that will likely never change, we would probably be stuck using only sentences and 'documents' (an article, a chapter of a book, etc).
So my questions, from a technical perspective, are:
1) do you find you get better results treating these longer texts as a series or words or a series of sentences, and should we be performing a double vectorization, were we vectorize all sentences, and then vectorize the longer text as a series of sentence vectors instead of using a word vectorization technique directly?
2) Assuming we have decided we want to vectorize sentences and longer sections of texts regardless of whether we use sentence vectors for the latter, do we use the same or different vectorization techniques/algorithms for both tasks, and do you have any recommendations on specific algorithms for either case?
Many thanks!
I'm quite new to the field myself. Many more experienced people, including the OP, I believe, would suggest that you use sentence transformers for this task. I personally don't understand how those are trained or fine-tuned, and I have never done it myself so far. What I do know is that sentence transformer results are horrible if you use them out of the box on anything other than the domain in which they were trained.
There's also the question of compute and RAM necessary to generate the embeddings. In one of my own projects, vectorizing 40M text snippets with sentence transformers would have taken 30 days on my desktop. So all I have worked with so far is fastText. It's several hundred times faster than any BERT model, and it can be trained quite easily, from scratch, on your own domain corpus. Again, this may be an outdated technique. But it's the one thing where I have some practical experience. And it does work quite well for creating a semantic search engine that is both fast and does not require a GPU.
The problem with fastText is that it only creates embeddings for words. You can use it to generate embeddings for a whole sentence or paragraph - but what it will do internally in this case is to just generate embeddings for all the words contained in the text, and then to average them. That doesn't give you really good representations of what the snippet is about, because common words like prepositions are given just as much weight as the actual keywords in the sentence.
Some smart people, however, figured out that you can pair BM25 with fastText: With BM25, you essentially create a word count dictionary on your corpus, that tells you which words are rare, and therefore especially meaningful, and which are commonplace. Then, you let fastText generate embeddings for all the words in your snippet. But, instead of averaging them with equal weights, you use the BM25 dictionary as your weights. Words that are rare and special thus get greater influence on the vector than commonplace words.
FastText understands no context, and it does not understand confirmation or negation. All it will do is find you the snippets that are using either the same words, or words that are often used in place of your query words within your corpus. But I find that is already a big improvement on mere keywords search, or even BM25, because it will find snippets that use different words, but talk about the same concept.
Since fastText is computationally "cheap", you can afford to split your documents into several overlapping sets of snippets: Whole paragraphs, 3 sentence, 2 sentence, 1 sentence, for instance. At query time, if two results overlap, you just display the one that ranks the highest.
Personally, I would imagine that doing document-level search wouldn't be very satisfying for the user. We're so used to Google finding us not only the page we're looking for, but also the position within the page where we can find the content that we're looking for. With a scientific article, it would be painful to having to scroll though the entire thing and skim it in order to know whether it actually is the answer to our query or not.
And with sentence-level, you'd be missing out on all context in which the sentence appears.
As the low hanging fruit, why not go with the units chosen by the authors of the articles, i.e. the paragraphs. Even if those vary wildly in length, it would be a starting point. If you then find that the results are too long, or too short, you can make adjustments. For snippets that are too short, you could "pull in" the next one. And for those that are too long, you could split them, perhaps even with overlap. I think for both the human end user and most NLP models, anything the length of a tweet, or around 200 characters, is about the sweet spot. If you can find a way to split your documents into such units, you'd probably do well, regardless of which technology you end up using.
You can also check out Weaviate. If you use Weaviate, you don't have to worry about creating embeddings at all. You just focus on how you split and structure your material. You could have, for instance, an index called "documents" and an index called "paragraphs". The former would contain things like publication date, name of the authors, etc.. And the latter would contain the text for each paragraph, along with the position in the document. Then, you can ask Weaviate to find you the paragraphs that are semantically closest to query XYZ. And to also tell you which articles they belong to.
You can also add negative "weights", i.e. search for paragraphs talking about "apple", but not in the context of "fruit" (i.e. if you're searching for Apple, the company).
These things work out of the box in Weaviate. Also, you can set up Weaviate using GloVe (basically the same as fastText), or you can set it up using Transformers. So you can try both approaches, without actually having to train or fine-tune the models yourself. With the transformers module, they also have a Q&A module, where it will actually search within your snippets and find you the substring which it thinks is the answer to your question.
I have successfully run Weaviate on a $50 DigitalOcean droplet for doing sub-second semantic queries on a corpus of 20+M text snippets. Only for getting the data into the system I had to use a more powerful server. But you can run the ingestion on a bigger server, and when it's done, you can just move things over to the smaller machine.
I can vaguely describe in a sentence the gist of an article I've read, or an image, and the proper result will usually be in the first page.
Of course, it doesn't always work, sometimes there are "hash collisions" so to speak, but I don't think the old algorithm would have been more successfully either, since if I knew the exact keywords to use, I wouldn't need to start with a vague description in the first place.
If not that, you get scam websites or other ad/malware infested trash.
It seems to do more correction for you, which is great if you're searching for common popular things. But any uncommon or precise query will often be misunderstood as something else.
Plenty of times, no matter how I reword my sentence or what sort of analogies I try to give it, I've had Google fail to give me something that I know exists and that I have to find some other way.
For the specific context of "I can find something I've already found", yes, it's useful. I just wish there was a way to change that context to "discovery mode" where it uses a different algorithm that is oriented toward finding new information. I want to find sites in the spirit of those old-fashioned sites that are minimally styled with dense information. And not just Wikipedia or a few "trusted" sources like it used to be in earlier times, but a more well-rounded result set.
I think the problem is that such sites are very difficult to find algorithmically, especially when it comes to their poor SEO. The reason they used to be so prevalent in the early 2000s search results is because that's mostly what the web was back then; a bunch of personal websites, blogs, etc.
To do that nowadays would require heavy (manual) curation, which obviously Google isn't interested in.
It especially noticeable to me that just in the last few months, Google has changed their algorithm so some product will be the first item on even the most generic search.
I don't even think it's that. The issue with older content isn't the search but the sorting. Google and to a lesser extent Bing (and thus DDG) have prioritized content claiming to be more recent relative to the time they were indexed in their display of results.
Showing more recent content is likely generally a safer bet for a search engine. More recent content is less likely to have suffered link rot and covers more recent developments in a subject.
Unfortunately Google et al don't really respect user preferences with respect to sorting. They have a financial incentive to show results they can turn into dollars (or cents).
I find it to be utterly terrible for that too.. even when I have verbatim strings from the thing I'm looking for it often simply doesn't show up. ... often because it rewrites the query into something about Kim Kardashian's butt and no amount of quotes or pluses will make it stop.
But sometimes it's a total miss when I want something very specific, and it just shows me other things I didn't ask for.
Also, Yandex is much better for reverse image search (similar images).
I can never do that. What I remember is so far from the wording used or how Google identifies the image that it never comes up. I end up scrolling back through my history trying to do it from the page title or domain
Whether one likes or hates current Google search results, their qualities and the changes from early search processes are clearly intentional and don't relate to how well Google does raw indexing.
[1] Docs: https://www.semi.technology/developers/weaviate/current/
[2] Github: https://github.com/semi-technologies/weaviate
[3] Wikipedia demo dataset: https://github.com/semi-technologies/semantic-search-through...
[4] Wikidata dataset: https://github.com/semi-technologies/biggraph-wikidata-searc...
Last week there was also a feature on Techcrunch about vector search and Weaviate: https://techcrunch.com/2021/12/11/2246180/
[1] Wikipedia Vector Search Demo with Weaviate: https://www.youtube.com/watch?v=IGB8vjCuay0
[2] Vector Search through Wikidata with Weaviate: https://www.youtube.com/watch?v=T4zlvknSbGc
[3] Demonstrations of Deep Learning: https://www.youtube.com/watch?v=5jbneytoKi0
[4] Weaviate's GraphQL API for Neurosymbolic Search: https://www.youtube.com/watch?v=K_2X48Tln9U
[5] Introducing the Weaviate Vector Search Engine: https://www.youtube.com/watch?v=AS_2U_INpKk
There is also this video about modern search engines and Weaviate on the AI Coffee Break YT channel: https://www.youtube.com/watch?v=YkK5IKgxp-c
Dumb question: How does Weaviate know that "Scandinavian" is close to "Finnish" ? The source not having "Scandinavian" at all. If their vectors are close, then the "vectorization" is quite standard for any text, and also per language?
EDIT: Just realized I didn't answer the second part of your question. Yes, the models are language-specific, but there are also multilingual models that work across a large no. of different languages.
[1] Sentence-BERT: https://sbert.net
[2] Weaviate Customizer with Out-of-the-box models: https://www.semi.technology/developers/weaviate/current/gett...
[1] https://www.pinecone.io/learn/
We are also actively researching the space, and just recently published a paper on improving Google's ScaNN: https://arxiv.org/abs/2112.02179
I read through your docs and figure that will be part of the approach.
An idea I had was to find similar, or "next best", cards for replacement in popular decks or to achieve similar effects in order to bring down the cost of EDH, Modern, etc. formats. I'm just getting back into the hobby again, so having a tool like this would make my wife and wallet happy :)
As for Pinecone itself, what are the main selling points as you see them for a simple application (e.g. comparing trigram-vectorized sets of strings) when compared to a home-rolled solution using postgres with array types? Better performance, ease of indexing, etc.?
In the meantime I can say moving to the dense vector + ANN search combo turns regular searches into semantic searches, which means more relevant results.
If that's the case for you, then you can use Pinecone to go further and make those results fast (<100ms), fresh (CRUD + live index updates), and filtered (apply single-stage metadata- filtering). All on a fully managed system that you can scale up/down with one API call.
(1) Pinecone uses dense vectors which can encode much more meaningful info, eg the actual 'semantic meaning' behind a sentence as we (people) would understand it, or the context in an image. Because of this, we can enable much richer, human-like interaction/search in your applications
(2) Performance wise, before joining Pinecone I was spending a lot of time with other dense vectors search tools like Faiss, and it isn't easy to get good or even reasonable accuracy and latency, particularly for large datasets. When I first used Pinecone, it took me maybe 10 minutes to figure everything out and start querying a reasonable dataset, search times were very fast and the accuracy incredible. Pinecone's tech is built by people that live and breath vector search, and what they've built outperforms anything I can build, even if I spend months trying to build it. I got better performance with Pinecone in 10 mins.
(3) Everything is production ready, no need to worry about deployment, security, maintenance etc, Pinecone deal with it and you can even use the service for free up to 1M vectors.
I'm dabbling in Postgres's full text search (ts_vector) for a small website, I know that is extremely simple compared to the offerings you provide, but your site has me quite interested in this space now.
Eager to learn more about this tech!
I believe Ant Financial has published an open source one but iirc the English language documentation is sparse.
I googled and did not find much pertaining to "ant financial" and "postgres". Perhaps your google-fu is better than mine...
It’s not a relational db, but it supports Graph-like connections between objects, which makes it really easy to model your relations.
It uses IVFFlat indexing, but could be extended to support product quantization / ScaNN.
I'm wondering if anyone has a suggestion of a SaaS or another alternative for my use case? Thanks!
E.g. searching for Huxley quote gives me silly blog posts about saving money.
Query: "The function of the brain and nervous system is to protect us from being overwhelmed and confused by this mass of largely useless and irrelevant knowledge, by shutting out most of what we should otherwise perceive or remember at any moment, and leaving only that very small and special selection which is likely to be practically useful."
Answer: "How to trick your brain into saving money"
Feel free to submit a new issues or merge request if you wish for new library added
(And how well does the technique of the article work wrt it?)
We've published a bunch of demo cases powered by vector database on GitHub. https://github.com/milvus-io/bootcamp
We have built Milvus vector database upon ANN libraries like faiss, annoy, nsmlib, etc.
We are aiming to create a cloud-scalable vector database. So Milvus comes to the crossroad of vector search and cloud database. There are many interesting system design topics in the development of Milvus 2.0. We will continue to share our experiences and thoughts on this topic.
A picture from a cartoon returns from logos to any type of drawing. A picture of a battery returns cars and shops. A picture of food worked as expected and I got more food pictures.
EDIT: ES integration PR: https://github.com/elastic/elasticsearch/issues/78473
Vespa also allows expressing hybrid sparse and dense retrieval (WAND for sparse, ANN via HNSW for dense). It's also easy to express multi stage retrieval ranking phases, as vector search alone is not achieving state-of-the-art ranking results, see https://blog.vespa.ai/pretrained-transformer-language-models...
Is there some people having feedback on this?
Not sure this helps but just mentioning in case.
Not sure how “free” it is though.
What is new here?
https://radimrehurek.com/gensim/auto_examples/tutorials/run_...
Or is it something different?
Everything old is new again! Again!