VectorDB: Vector Database Built by Kagi Search
vectordb.com
vectordb.com
Here is an example Colab notebook [1] where this is used to filter the content of the massive Kagi Small Web [2] RSS feed based on stated user interests:
[1] https://colab.research.google.com/drive/1pecKGCCru_Jvx7v0WRN...
Edit: I see you run a vectordb company, so your question makes more sense
We generally avoid using embedding services to beginwith.. outside calls is the special case. Imagine something heavy like video transcription via Google APIs, not the typical one of 'just' text embedding. The actual embedding is generally one step of broader data wrangling, so there needs to be a good reason for doing something heavy and outside our control... Which has been rare.
Doing in the DB tier is nice for tiny projects, simplifying occasional business logic, etc, but generally it's not a big deal for us to run encode(str) when building a DB query.
Where DB embedding support gets more interesting to me here is layering on additional representation changes on top, like IVF+PQ... But that can be done after afaict? (And supporting raw vectors generically here, vs having to align our python & model deps to our DB's, is a big feature)
Their simple suggestion to extend to longer texts is to pool/avg sentence embeddings - but I'm not so sure that I want that; for instance eg that implies order between sentences didn't matter. If I were forced to use sentence transformers for my use case, then the real fix would be to train an actual pooling model atop of the sentence embedder, but I didn't want to do that either. At that point I stopped looking into it, but I'm certain there are newer models out nowadays that have both better encoders and handle much longer texts. The one nice thing about the sentence transformer models though is that they are much more lightweight than eg a 7B param language model
> Thanks to its low latency and small memory footprint, VectorDB is used to power AI features inside Kagi Search.
Is there any limitation to this? Does it work with text 500-1000 words? Does it work well with text that aren’t full sentences? Ie. Are just a collection of phrases?
https://github.com/kagisearch/vectordb/blob/main/vectordb/st...
Not a knock against this project, I see where it can be helpful.
If you really didn't want PyTorch/Transformers, you could consider exporting your models to ONNX (https://github.com/microsoft/onnxruntime).
At least for use cases where there are clusters of many similarly formatted documents, it would be cool to have a way of easily customizing chunking.
[0] https://help.kagi.com/orion/company/hiring-kagi.html#full-ti...
[1] CrystalConf talk from a Kagi Tech Leas https://www.youtube.com/watch?v=r7t9xPajjTM
not sure if there is something similar in Python
wish there is something simpler but statically-compiled with python syntax, something like crystal but for python developers.
https://docs.scala-lang.org/scala3/reference/other-new-featu...
Looks like it uses one of these, depending on your settings:
Fast model: google/universal-sentence-encoder/4
Multilingual model: universal-sentence-encoder-multilingual-large/3
Normal model (Alternative): BAAI/bge-small-en-v1.5
Best model: BAAI/bge-base-en-v1.5
If the typical user will use the thing once and then forget about it or if they are not going to keep it in mind very often make sure to give it a descriptive name.
But for some things having a descriptive name might not be the most important aspect of its name. Having a findable or memorable name might be more important. And maybe sometimes giving it a descriptive name might actually become a problem because its function might change over time.
A british recruiter was quite insistent to have a call… turns out because she read "informatica" on my profile (as in, laurea in informatica), asked me how many years of experience I had with the tool, I replied "I just heard of it 2 minutes ago when you first mentioned it". Then got mad at me for having written the word "informatica" on my profile.
Something like “Kagi vector search for databases” at least doesn’t leave anything up for misinterpretation.
You are not wrong about the performance from Rust, but LanceDB is inherently written with performance in mind. SIMD support for both x86 and ARM, and an underlying vector storage approach that's built for speed (Lance)
Not necessarily a good thing when the product is made by a VC backed startup that may die or pivot in six months leaving you the need to maintain it yourself.
You can write performant code in any language. For example, for standard keyword search, I wrote a component to make sparse/keyword search just as efficient as Apache Lucene in Python. https://neuml.hashnode.dev/building-an-efficient-sparse-keyw....
I referenced this article below but will reference it again here too. https://neuml.hashnode.dev/building-an-efficient-sparse-keyw....
You can write performant code in any language if you try.
PyTorch, quoting themselves, is a Python binding into a monolithic C++ framework; also optionally depending on existing libs like mkl etc.
> You can write performant code in any language if you try.
Unfortunately, only to a certain extent. Sure, if you just need to multiply a handful of matrices and you want your blas ops to be blas'ed where the sheer size of data outweighs any of your actual code, it doesn't really matter. Once you need to implement lower-level logic, ie traversing and processing the data in some custom way, especially without eating extra memory, you're out of luck with Python/numpy and the rest.
I guess this is a pretty legitimate take, but in that case VectorDB looks like (from the got repo) it makes huge use of libraries like pytorch and numpy.
If numpy is fast but "doesn't count" because the operations aren't happening in python, then I guess VectorDB isn't in python either by that logic?
On the other hand, if it is in Python despite shipping operations out to C/C++ code, then I guess numpy shows that can be an effective approach?
LAPACK and other ancillary stuff could be Fortran or C.
Anyway, every language calls out to functions and runtimes, and compiles (or jits or whatever) down to lower level languages. I think it is just not that productive to attribute performance to particular languages. Numpy calls BLAS and LAPACK code, sure, but the flexibility of Python also provides a lot of value.
How does Numba fit into this hierarchy?
Yes, pure Python is slower and takes up more memory. But that doesn't mean it can't be productive and performant using these types of strategies to speed up where necessary.
> Where do you draw the line?
Drawing the line at native python, not pulling in packages that are written in another language. Packages written in python only are acceptable in this argument.
> But that doesn't mean it can't be productive and performant using these types of strategies to speed up where necessary.
No one said it couldn't. What we're saying is that it pure python is 'slow' and you need to escape from pure python to get the speedups.
https://stratoflow.com/efficient-and-environment-friendly-pr...
Ok, I have to call this statement out. Mypy was released 15 years ago, so Python has had optional static typing for as long as you've been programming in it, and you don't know about it?
I guess it's going to take another fifteen years for this 2008 trope to die.
Re-reading my comment, no, I did not. I said it has static typing.
I used Python type hints and MyPy since long before I used TypeScript, and I have to say that TypeScript's take on types is just plain better (that doesn't mean it's good though).
1. More TypeScript packages are properly typed thanks to DefinitelyTyped. Some Python packages such as Numpy could not be properly typed last I checked, I think it might change with 3.11 though. Packages such as OpenCV didn't have any types last I checked.
2. TypeScript's type system is more complete, with better support for generics, this might change with 3.11/3.12 though.
3. TypeScript has more powerful type system than most languages, as it is Turing-complete and similar in functionality to a purely functional language (this could also be a con)
Why? That was just mean for no reason!
But, everybody else seems to agree so maybe I’ve been had.
This repo is basically just a nice api and the needed chunking and batching logic. Using lancedb, you'd still have to write that, as exemplified here: https://github.com/prrao87/lancedb-study/blob/main/lancedb/i...
When I use Common Lisp or Racket I roll my own simple vector embeddings data store, but that is just me having fun.
We needed a low latency, on premise solution that we can run on edge nodes with sane defaults that anyone in the team can whim in a sec. Also worth noting is that our use case is end to end retrieval of usually few hundred to few thousand chunks of text (for example in Kagi Assistant research mode) that need to be processed once at run time with minimal latency.
Result is this. We periodically benchmark the performance of different embeddings to ensure best defaults:
https://github.com/kagisearch/vectordb#embeddings-performanc...
No one has mentioned wallabag yet, so wanted to. Been working well for me - has apps and extensions. If you’re not excited to self-host - https://www.wallabag.it/en has been flawless with the exorbitant price of… 11 euro a year.