What is a Vector Database? (2021)
pinecone.io
pinecone.io
Supabase wrote a solid tutorial[1] (you don't need to run it on Supabase).
0 - https://github.com/pgvector/pgvector
1 - https://supabase.com/blog/openai-embeddings-postgres-vector
Current concerns are the scaling and recall performance.
The author is looking at product quantization along with other ideas: https://github.com/pgvector/pgvector/issues/27
More details on product quantization: https://mccormickml.com/2017/10/13/product-quantizer-tutoria...
A nice repo that tracks the ANN relative performance of different indexes: https://mccormickml.com/2017/10/13/product-quantizer-tutoria...
Also shoutout to Weaviate because they have great docs, are open source and have very informative YouTube channel.
We're building an AI data analyst. You can ask questions of your database and get answers immediately. We also auto generate entire dashboards based on common patterns (e.g. a "Sales Dashboard", "Marketing Dashboard", "Finance / Burn" etc.).
If you want to give it a try (there's a demo database embedded in the app), you can use it here: https://ui.definite.app/
Not that document databases don’t have their place, but…MongoDB is webscale and all that.
I ended up choosing Weaviate specifically because of the nice docs, but beyond that, time will tell.
https://weaviate.io/developers/weaviate/api/graphql/filters
https://weaviate.io/blog/hybrid-search-explained
I have a ChatGPT session where I have asked it to do a hybrid search using filtering, pg fts and vector search. Looks reasonable just need to test it and write it up somewhere.
It seems like if the goal is to "play around with vector databases", why not just install it on your local machine? Part of using these tools is learning how they work and configuring them yourself.
If the goal is "start developing products using vector data bases" then it seems like you would surely want something a bit more under your control than using replit.
Thanks for building this.
disclaimer: i'm a founder at algora.io, the platform that enables these paid contributions
https://docs.algora.io/bounties/payments#compliance
If algora.io didn't charge %23 of the bounty I would have tried to contribute. It felt unfair to me.
If they weren't paying Algora to do it, they'd be paying their own staff to do it. Either way, the extra percentage wouldn't be part of the bounty.
I'd gladly pay 23% just to not have to worry about the logistics of bounty payments.
@qdrant_team: perhaps you should look into offering it as a service, a la pinecone.
edit: oops just checked your (updated) website and notice you have an offering already. Congrats! will check it out. ty =)
A lot of people worry about copy-cat services, but it's kind of rare that someone will be able to compete with you as the original in hosting your own service as well as you can. Especially when you consider support and maintenance requirements of a new product you aren't personally developing.
I could see copy-cat services being more of an issue in the late stage of a product though? When everyone knows lots about how to stand it up and use it?
[1] https://github.com/coralproject/talk/blob/develop/LICENSE
The concern isn't random small companies. The concern is the big cloud providers like AWS, Azure and Google. And you are right, they aren't going build out a hosted version of your product until there is enough traction. But at that point, customers might indeed trust them more than you to run your own software! Redis and Elastic ran into this problem for example.
The most likely scenario though - is never getting traction - so anything to improve traction such as permissive licensing is probably a better tradeoff.
Thanks for pointing it out so folks can be wary of AWS (and similar) eating their lunch like they have countless other SaaS services!
About the integrated vector search: https://docs.cozodb.org/en/latest/releases/v0.6.html
It also does duplicate detection (Minhash-LSH) and full-text search within the query language itself: https://docs.cozodb.org/en/latest/releases/v0.7.html
HN discussion a few days ago: https://news.ycombinator.com/item?id=35641164
Disclaimer: I wrote it.
None of these explains what an embedding really is. My best guess is that the embedding represents the meaning of the natural language string it was generated from, such that strings with "similar" embeddings have similar meaning. But that's just speculation.
You use a machine learning model (like word2vec, OpenAI, etc.) to produce an "embedding" that describes the image, text, video, etc., which is your "vector".
For all of the other images in your database, you also run them through the same model, and store their embedding vectors in the vector database.
Then, you ask the database "I have this vector, what are the most similar vectors, and what are their primary keys, so I can see what content they refer to".
Think: you want to implement google "search by image". This is the basics of how you'd do that.
Why use the word "embedding" if there are already much more familiar words for it (isn't this the same as feature vector)?
I want to convince myself that this isn't similar to blockchain. In the sense that blockchain renamed an old and simple idea and advertised it as something complex and groundbreaking...
Also, relational databases or graph databases have a reach theory that results in many interesting sub-problems, each interesting in its own right, to contrast this with "document databases", which have no theory, and nothing interesting behind it. So, if I were to invest my time learning about one w/o a financial incentive to do so, I'd not want to concentrate on some accidental concept that just happened to solve an immediate problem, but isn't applicable / transferable to other problems.
For example, graph databases and relational databases create interesting storage problems wrt' optimal layout for various database components. If hash-table is all there is to the vector database, then it's not an interesting storage problem.
Similarly, with querying the database: if key lookup from a hash-table is all there is, then it's not an interesting problem.
Yeah, you've got it. A mapping from words to vectors such that semantic similarity between words is reflected in mathematical similarity between vectors.
An idea of how you might train this thing: lets say the words "king" and "queen" are being embedded. In your training data there are lots of examples where "king" and "queen" are interchangeable, for example in the sentence "The ___ is dead, long live the ____", either word is appropriate in either slot, so each time we see an example like this we nudge "king" and "queen" a little closer together in some sense. However you also find phrases where they are not interchangeable, such as "The first born male will one day be ____". So when you see those examples you nudge "king" a little closer in some sense to other words which appropriately complete the sentence (which does not include "queen" in this case).
In this way, repeated over a giant training set with thousands of words, concepts like "male/female" and "royalty", "person/object" and tons of others end up getting reflected in the relationships between the vectors.
These vectors are then useful representations of words to ML models.
As for the number of dimensions, in a sense they are a training variable just as the content itself. The more dimensions you utilize for your embeddings the more complex your relations can be during clustering. Too many dimensions can easily lead to over fitting however and too little dimensions can usually not accurately represent the training corpus.
If the inputs are images, you may find that some dimension scores e.g. how much blue there is in the image. Though often it's not that simple (there could be multiple dimensions that relate to how blue the image is, especially if the embedding dimensionality is large, which it does tend to be these days. Though you could reduce the embedding dimensionality first using PCA, and see what input images correspond to high/low values of the first principal component, etc.).
Starting with: what do you store in it?
Maybe sentence/vector pairs. But what does that give you? What do you do with that data algorithmically? What's the equivalent of a SELECT statement? What's the application that benefits an end user? That part still seems rather hazy.
You measure the cosine distance between documents, or between search queries and documents. (Cosine is fast, there are other distance metrics).
The vector database queries will do things like given one embedding (document or query) find the nearest embeddings (documents). Or given two embeddings (e.g. a query and a context) with a weight for each one, find the ones that triangulate to being near both.
An autoencoder is a model that takes a high dimensional input, distills it down to a low dimensional middle layer, and then tries to rebuild the high dimensional input again. You train the model to minimize reconstruction error, and the point is then that you can run an input on just the first half to get a low-dimensional representation that captures the "essence" of the thing (in the "latent space"). In this representation, images that are similar should have similar "essences", so their latent vectors should be near to each other.
The low dimensional representation must do a good job capturing the "essence" of your things, otherwise your reconstruction error would be large. The lower the dimension you manage to use while still managing to reconstruct your things, the better of a job it must do at making those parameters really encode the salient features of your thing without wasting any information. So similar things should be encoded similarly.
So imagine you've got a database of images, and you have a table of all of the low dimensional encoded vectors. You want to do a reverse image search. The user sends you an image, you run the encoder on it to get the latent representation, and then you want to essentially run "SELECT ei.image_id FROM encoded_images ei ORDER BY distance(encode(input_image), ei.encoding) LIMIT 10".
So you want a database that supports indexes that let you efficiently run vector similarity queries/nearest neighbor search, i.e. that support an efficient "ORDER BY distance(_, indexed_column)". Since the whole process was fuzzy anyway, you may actually want to support an approximate "ORDER BY distance" for speed.
In practice apparently the encoding might be taking the output of the first or nth layer in a deep network or something rather than specifically using an autoencoder. Or you may have some other way to hash/encode things to produce a latent representation that you want to do distance searches on. And of course images could instead be documents or whatever you want to run similarity searches on.
In that case you might store vectors representing a user based on some features youve selected, or a word embedding of their common genres/tags.
To actually search this thing, you need something to compare against. You could directly use the word embeddings of the search query. You could also do a search against your existing method, and then use the top results from that as a seed to search your vectors.
Since everything's a vector, you can also ask questions like "what musician is similar to Tom AND Sally" by looking for vectors near T+S. T-S could represent like Tom but not like Sally, etc.
So the answer to what do you store is, what will be your seed to search against?
Perhaps there is also some unitary vector operation which directly corresponds to negation (not Sally)? Perhaps multiplying the vector by -1? Or would (-1)S rather pick out "the opposite of Sally in conceptual space" instead of "not / anyone but, Sally"? And what about logical disjunction (union)? One could go further here, and ask whether there is an analog to logical quantifiers. Then there is of course the question whether there is anything in vector logic which would correspond to relations, binary predicates, like R(x, y), not just unitary ones, etc.
(Sorry for rambling, I'm thinking out loud here.)
So, coming myself from a database background but working in search, the SELECT statement (and joins) probably aren't the best way to get your head wrapped around things. I would think of the vector as a unique key for a record, and only using a LIKE statement for all my queries, but one that will return a probability of a match instead of an actual match.
A great use case is to think about similarity, where we want the things that are closest to what we want to see, but there isn't an exact match.
For example; a user gives me a sentence that says, "How long do I have to be with the company before I get a 401K match?". My vector store has a bunch of vectors including "A new employee will be eligible for 401K after 6 months." ,and, "The 401K program is run by <MEGACORP X>."
I would like to be able to see that the first vector is a closer match to the user sentence than the second, and by how much. I would also like to do this without having to change my code much based on the structure of the text. Luckily, there is a very simple algorithm for doing this (cosine similarity) that doesn't change regardless of the sentence structure or the question answered. Also, it doesn't matter what kind of question/answer you do as long as it can be vectorized, so you could even give me a vector representing an image and I can give you an image that is most similar.
Here is the most interesting thing about vectors -- with very little effort they turn the english language into a programming language.
Instead of typing "SELECT document_id, document_name, document_body FROM documents WHERE (document_body LIKE '%401K%' AND document_body LIKE '%match%' AND document_body LIKE '%existing employee%') FROM documents" I can just ask, "How long do I have to be with the company before I get a 401K match?" and I will get back a result and a match probability. How I change my text will change the matches, and can do so in ways that are profound and unexpected. Note that the SQL query I gave would not return any values because I didn't have any documents that had the term "existing" in them. Building the correct SQL query could be quite complex it comparison to just using the text.
This is pretty great for long-tailed search, q&a, image search, recommendations, classification, etc.
BTW, I am biased, I work for Elastic (makers of Elasticsearch) and we have been doing traditional search forever, and vector/hybrid search for the last few years.
For example, a neural net model accepts a massive number of input values that directly map to the input. So those initial values don't add any info. But a layer further inside the model, with fewer values and probably close to the end, is smaller and should reflect what the model's learned. Like a lot of deep learning, three values work but don't give much insight.
If I'm wrong, I hope somebody more knowledge corrects me. I got my understanding from basic into tutorials and Wolfram's essay on ChatGPT: https://writings.stephenwolfram.com/2023/02/what-is-chatgpt-...
Each vector is an array of n floats that represent a location of a thing in an n-dimensional space. The idea of learning an embedding is that you have some learning process that will put items that are similar into similar parts of that vector space.
The vectors don’t necessarily need to represent words and the model that produces them doesn’t necessarily to be a language model.
For example, embeddings are widely used to generate recommendations. Say you have a dataset of users clicking on products on a website. You could assume that products that get clicked in the same session are probably similar and use that dataset to learn an embedding for products. This would give you vector representing each product. When you want to generate recommendations for a product, you take the vector for that product and then search through the set of all product vectors to find those that are closest to it in the vector space.
The general idea is that the items you are embedding may vary in very many different ways, so trying to map them into a low dimensional space based on similarity isn't going to be able to capture all of that (e.g. if you wanted to represent faces in a 2-D space, you could only use 2 similarity measures such as eye and skin color). However a high enough dimensional space is able to represent many more axis of similarity.
Embeddings are learnt from examples, with the learning algorithm trying to map items that are similar to be close together in the embedding space, and items that are dissimilar to be distant from each other. For example, one could generate an embedding of face photos based on visual similarity by training it with many photos of each of a large number of people, and have the embedding learn to group all photos of the same person to be close together, and further away from those of other individuals. If you now had a new photo and wanted to know who it is (or who it most looks like), you'd generate the embedding for the new photo and determine what other photos it is close to in the embedding space.
Another example would be to create an embedding of words, trying to capture the meanings of words. The common way to to this is to take advantage of the fact that words are largely defined by use/context, so you can take a lot of texts and embed the constituent words such that words that are physically close together in the text are close together in the embedding space. This works surprisingly well, and words that end up close together in the embedding space can be seen to be related in terms of meaning.
Word embeddings are useful as an input to machine learning models/algorithms where you want the model to "understand" the words, and so it is useful if words with similar meaning have similar representations (i.e. their embeddings are close together), and vice versa.
Imagine you wanted to go from words to numbers (which are easier to work with mathematically), like you wanted to assign a number to some words.
How could you do it? Well you could do it randomly: cat could be 2, dog could be 10, sweater could be 4.534 and frog could be 8.
Not super useful, but hey - words are now numbers! How can we make this "better"?
What if we decided on a way to put words on a line - let's say we ordered words by how much they had to do with animals. Let's say 10 meant it's a very animal-related word, and 0 is very not-animal related. So cat and dog would be 10, and maybe zoo would be 9, and fur could be 8. But something like sweater would be 1 (depending if the sweater was made from animal wool...?)
What now? Well what's cool is that if you assign words on that "animal-ness" line, you can find the words that are "similar" by looking at the numbers that are close. So, words whose value is around 6 are probably similar in meaning. At least, in terms of how much they relate to animals.
That's the core idea. Ordering words by animal-ness is not that useful in the real world, so maybe we can place words on a 2d grid instead of a line. Horizontally, it would go from 0 to 10 (not animal at all - very animal) and vertically, it could be ordered by brightness - 0 for dark, and 10 for bright.
So now, bright animals will congregate together in one part of the grid, and dark non animals will also live close together. For example, a dark frog might be in the bottom right at position (10, 0) - very animal (right end of the x axis) but not bright (bottom of the y axis). Any other word whose position is close to (10, 0) would presumably also be animal-y and dark.
That's really it. The magic is that... this works in thousands of dimensions. Each dimension being some way that "AIs" see words / our world. It's harder to think about what each dimension "is" or represents. But embeddings are really just that - the position in a space with a huge number of dimensions. Just like dark frogs were (10, 0) in our simple example, the word "frog" might be (0.124, 0.51251, 0.61, 0.2362, 0.236236, ..............) as an embedding.
That's it!
The example you used going from 1 to 2 to n dimensions really made sense
I get the why the techniques are suitable, but I just assumed who ever wants to do this kind of retrieval can probably implement a suitable Approx. NN library themselves?
Especially so, because getting good embeddings is the hard part, not the search?
Searching efficiently is a problem, and there's several open source and proprietary solutions but I don't get how you can put it in the "everyone should roll their own" category.
There are plenty of options available to run your own local vector database, txtai is one of them. Ultimately depends if you have a sizable development team or not. But saying it is impossible is a step too far.
1B vectors * 300dimensions * float32 (4 Bytes) ~= 1.2TB
This pretty much still runs on consumer hardware.
Just run that on a 4TB nvme ssd, or a RAID array of ssd's if you're frisky.
Sure, role your own, but don’t act like making a highly scalable database is a weekend project.
> You do realize you have to query an index of all of that data for every single query your use makes right? Computing that index is not entirely trivial, nor is the operation of partitioning the data so it fits in ram across a pool of nodes.
I don't know what any of this means -- and it sounds like you're slapping a bunch of terminology together, rather than communicating a well-thought-out idea.
Yes, in the general case you're going to have to use an index. Computing an index or a key to that index? Computing the index is a solved problem, that does not have a hard real-time component -- you can do it outside of normal query executions. Computing the key to the index on each query is also a solved problem.
Have dimensions stored in columnar format, generate a sparse primary index on said columns, and then use binary search to quickly find the blocks of interest to do a sequential search on viz. distance function. Or you could even just use regular old SS trees, SR trees, or M Trees for high-dimensional indexing -- they're not expensive to use at all.
There, you can easily run a query on a single dimension (1 billion entries) under a second. You want 300 dimensions? Ok, parallelize it. 128 threads, easy. At most this will take 3 seconds if everything is configured properly (big IF, that seems like few can get right).
This is literally a weekend project. Anyone can build something like this, but not everyone has the integrity to be upfront about how they're reinventing the wheel, and spinning it like they've just broken ground in database R&D.
All of it fits on a single machine on one or a few big, fast SSDs.
So that kind of dataset fits on a small SSD. :-)
While Pinecone isn't available as a self-hosted option (see many comments with alternatives), we do offer the option of running Pinecone for you on a managed VPC, and we do have SOC2 compliance, and we do pass enterprise-level security reviews regularly. Whether that's sufficient is up to you of course.
First time here? Just kidding. But not.
You have to separate the VC hype with the product, because the VCs always need something to overhype. Half these people were pumping money in to crypto and whatever-the-hell-web3-is/was just a couple months ago, this is just the next thing they like. Half these companies probably aren't remotely good companies.
The VC money hardly ever makes sense.
There are plenty of LLMs to choose from with regard to finding sources of embeddings. Some free, some for money.
struggling to see the reason for the sudden demand
It is also effectively "roll your own Google/Shazam/whatever", which probably makes for a fancy demo to those who don't know how trivial it is to implement.
Basically investors are morons on average.
I don't see it addressed in the article, but Elastic 8 has ANN support, and every other feature you'd expect out of a ranking system. Vectors are only one piece of the puzzle for building such a system. (honest question, not trying to troll, as I truly do <3 these pinecone articles)
(Similarly, Y not Solr, Vespa, etc etc) :)
We always encourage folks to do their own testing. Everyone has different performance requirements, data shapes/sizes, budgets, and expectations of the user experience.
Elasticsearch is a great option. But clearly there's a large cohort of smart teams that decided the combination of performance + cost + scale + [etc] on Pinecone makes more sense for them.
IMO - the real reason "Y Not Elasticsearch" is not because they're dumb or its bad. It's actually because they're not building for the search / AI market like you all are :)
When someone runs out of RAM with their Numpy array, they google, and you guys come up really speaking to that audience, building features, showing people how to build specific solutions, etc.
DR is taking a long embedding and doing something to make it shorter. An easy to follow method for this is minhash.
Neighborhoods is representing a cluster of embeddings with a single representative to speed up comparisons. For example, find me the two closest representatives then doing a deeper comparison on all the residents.
Now the feature I haven't seem that will probably cause me to build instead of buy. Most seem designed for a single organization and a single use. For example, Spotify song recommender.
I would like to store embedding from multiple models and be able to search per model. I would also like fine grain user access control, so users could search their embeddings and grant access to others.
If you mean "user access control" within your company, there are basic access controls within Pinecone. See: https://docs.pinecone.io/docs/add-users-to-projects-and-orga...
If you mean for your end-users, you can use namespaces again to separate embeddings for different users inside one index. See: https://docs.pinecone.io/docs/multitenancy
There isn't yet a combination of the two, where you provide Pinecone API access to end-users.
[1] https://gpt3experiments.substack.com/p/building-a-new-embedd...
I can't speak for the competition, but weaviate seems to support them: https://weaviate.io/developers/weaviate/concepts/binary-pass...
More so, I am wondering how it is that Pinecone manages to land on the HN front page so frequently given the large number of alternatives (see all the other comments in this thread). It suggests to me a coordinated marketing effort (brigading, cronyism, etc.)
Oh man, now I sound like I'm part of the brigading :) FWIW I recently evaluated pinecone and alternatives and decided not go with pinecone (pgvector).
[1] https://www.pinecone.io/learn/
[2] https://docs.pinecone.io/docs/examples
Having myself recently bootstrapped an understanding of language models et al, I would not be surprised if the pinecone learning center gets a lot of traffic.
for the scale you are saying "larger scale": At the few million documents scale I would just suggest using just any libary, eg. `hnsw` in `nmslib` or `faiss`.
I just did some benchmarks with 1M docs, `cosinesimil_sparse` on `78628` dimensional binary vectors (nmslib `hnsw`) -> 30 seconds to build the index, and can process a batch of 100 document query in 3ms (Each with 100 KNN). Based on this question, i just put a loop over it and it handled 1000 random queries (non batched) in 1.11 seconds. (~1 GB peak RAM usage, and using 24 threads)
All in all, my personal opinion is: even up to few "millions" scale, i'm finding using the underlying libraries (`faiss` and `nmslib`) significantly easier than using the wrapper tools / databases (milvus and pinecone). I don't really get the point of a separate piece of infra for something that is essentially ~15 lines of python at most scales that matter (~few millions). (Note, in the ~10k-100k scale or less, simple numpy and sort seems to be fast enough (and exact) or just exact NN w/ sklearn.neighbors)... And when you push to scales that it does start breaking (100 million+), then the database versions seem to break as well (and require fiddling with lots of bespoke config)
Currently I use Elasticsearch with the Open Distro approximate kNN plugin by the way.
(I can't speak to Milvus 2.x as we are probably not going to upgrade to that for a number of non-performance reasons)
I remember reading about how Google maps did a very similar thing to figure out which points of interest to load based on your coordinates and zoom.
Can't we repurpose that technology? Or did those bake in assumptions around being 2 dimensional (while this is highly dimensional?)
Once this is done, the search heuristics are not difficult (find the cells to explore and return nearest neighbors).
Orthogonality is expected, for example. Proximity is really rare. See: https://softwaredoug.com/blog/2023/02/28/probability-of-dot-...
E.g., KNN is very fast, if and only if you have a high performing R-tree query.
so while the 2d algos might "fall out" of the sophisticated cases, the sophisticated cases will need to go in a different optimization direction.
But as the other comments have mostly said, it's mainly dimensionality and scale differences that drive the design differences (e.g. graphs end up working better than trees in high dimensions)
What are vector DBs used for? > Storing and search through embeddings at scale, which are created and consumed by LLMs and other AI models for applications like semantic search and chatbots (eg, to avoid hallucinations).
Why use a managed vector DB like Pinecone instead of [Faiss, pgvector, self-hosted thing, numpy.array, etc]? > Usually comes down to scale and convenience. If you're dealing with a small amount of embeddings, say anything less than 10M, you're probably fine just reaching for the closest and most convenient option. (We try to make Pinecone that convenient option, and our free plan holds up to ~100k 1536-dimension embeddings.) If you're dealing with larger scale -- say hundreds of millions to billions of embeddings -- and have strict performance requirements, and aren't thrilled by the thought of managing your own vector database like we are, then you should consider Pinecone. It turns out there's a sufficiently large population that falls into the latter category, just as with any other database category.
If you're new to this the best place to "see for yourself" is our free plan (https://app.pinecone.io) and collection of examples (https://docs.pinecone.io/docs/examples).
Allow me one more plug: We're hosting a webinar next week about testing Pinecone performance with your own data and performance requirements. I have a feeling lots of folks reading this would find that useful. → https://pinecone-io.zoom.us/webinar/register/WN_z9JqLjLGTyu4...
Plug: If you're ever looking for an open source alternative to Pinecone, we recently added vector search to Typesense: https://typesense.org/docs/0.24.1/api/vector-search.html
The key thing is that it's in-memory and allows you to combine attribute-based filtering, together with HNSW-based nearest-neighbor search.
We're also working on a way to automatically generate embeddings from within Typesense using any ML models of your choice.
So Algolia + Pinecone + Open Source + Self-Hostable with a cloud hosted option = Typesense
I wouldn't call postgres lightweight of course, but it's definitely lightweight in the sense that it doesn't add a whole bunch of new garbage to an otherwise traditional application.
In an image, each pixel is a dimension but it does not have any meaning in itself - you need to look at the rest of the image to understand that the pixel is part of a cat. Embeddings is a way to represent this meaning. Think about it as a “summary” of a image / document.
So... a vector database can be organized to quickly retrieve objects with particular characteristics, without rigidly defining what those characteristics are.
Have I got it?
What does "near-perfect" mean in this context, type 1 or type 2?
Searching for similar vectors is basically the (approximate) KNN problem, although I imagine more specialized search methods might apply depending on what you are doing.
Recently published an article discussing this: https://neuml.hashnode.dev/customize-your-own-embeddings-dat...
Elasticsearch uses HNSW, not sure what options they have but quantization/compression will help reduce disk storage requirements. Alternatively, you can look at dimensionality reduction algorithms and only store that output in ES. Or pick a model with a small number of dimensions. For example https://huggingface.co/sentence-transformers/all-MiniLM-L6-v... only has 384 dims vs 768/1024/2048/4096.
The context was a very underwhelming side project of his: A movie search engine but you had to use the exact titles of the movies to get results. It only revealed that he doesn't appreciate what similarity search actually is.
It feels almost blasphemous to call a Karpathy side project underwhelming. He is a genius and it really felt unlike him to write that "just use np.array" tweet.
Surely the whole point of a vector "database" in that context would be to store semantic sentence embeddings of the movies titles to support approximate / semantically-related search ? Could do the same thing for movie plot synopsis too - allow user to search via vague descriptions of movie. ChatGPT actually does very well at this, although massive overkill.
But when you start to require filtering or combining the vector search with a lexical search, then something like Pinecone, Vespa, Qdrant, Lucene-based options (e.g. Solr and ES) etc. become a lot more practical than you building all that functionality yourself.
For NLP use-cases, you can use a vector database to index the embeddings of your texts.
For exemple, if you implement a document retrieval (a search engine like Google), you train a transformer model that takes a text as input (the content of your document) and produces a vector of number as output (the embedding). You then index your documents by transorming them to their embeddings and storing them inside your vector database.
When you want to perform a query using keywords, you transform your keywords into a vector, and then ask your vector database to send you the most similar documents, using a similary function such as the cosine function.
( aside from all the per vector associated meta-data, various dimensional reduction, nearest neighbour, etc, operations describe in the article ).
The question comes from is it the CompSci | HN | AI domain nomenclature to assume a vector database is made up of vectors that are all of the same dimension over a continuum (eg. N real numbers for fixed N) or are vector databases made up of mixed vectors (no fixed dimension) and discrete values, etc.
I ask as the linked article doesn't specify but does appear to imply.
You would not mix and match embedding models (e.g. with differing dimensions) at look-up time. The target vector table assumes you will look it up with a vector created from the exact same embedding model and version that was used to backfill it.
The API documentation for a look up operation may be more illuminating here:
>vector (array of floats)
>The query vector. This should be the same length as the dimension of the index being queried. Each query() request can contain only one of the parameters id or vector.
In any case your description is strangely reductionist. The important thing about a vector db is that is typically designed to store embeddings used in various ML applications. So say you are doing NLP you can tokenize some input and then store the token and positional embeddings in a vector db and then use it for similarity search, training etc.
> your description is strangely reductionist.
Sure - pure | applied math background, old enough to have used Postgres when it was known as Ingres, to have patched in Spatial relations before it had the GIS functions it has now, and to have written libraries for GIS linked { 256 | 1024 | 2048 } D vector databases for signal aquisition | processing.
I'm late to the 'modern' discussions & just checking my read - I can think of applications for mixed dimensions and discrete space vectors and there are analogs to trad R^N ops for those cases.