Serverless Semantic Search, Free tier only
qdrant.tech
qdrant.tech
1. Install sentence-transformers [1]
2. Initialize the MiniLM model - `model = SentenceTransformer('all-MiniLM-L6-v2')`
3. Embed your corpus [2]
4. Embed your queries, then search the corpus
This runs on CPU (~750 sentences per second), and GPU (18k sentences per second). You can use paragraphs instead of sentences if you need more text. The embeddings are accurate [3] and only 384 dimensions, so they're space-efficient [4].Here's how to handle persistence. I recommend starting with the simplest strategy, and only getting more complex if you need higher performance:
- Just save the embedding tensors to disk, and load them if you need them later.
- Use Faiss to store the embeddings (it will use an index to retrieve them faster) [5]
- Use pgvector, an extension for postgres that stores embeddings
- If you really need it, use something like qdrant/weaviate/pinecone, etc.
This setup is much simpler and cheaper than using a ton of cloud services to do embeddings. I don't know why people make semantic search so complex.I've used it for https://www.endless.academy, and https://www.dataquest.io and it's worked well in production.
[2] https://www.sbert.net/examples/applications/semantic-search/...
[3] https://huggingface.co/blog/mteb
[4] https://medium.com/@nils_reimers/openai-gpt-3-text-embedding...
But the general dialogue around AI-related tools is surprising to me. The production parts of the langchain, embeddings, etc tools can usually be built in a few hours with better observability, performance, and maintainability.
https://github.com/azayarni is a contributor (andre-z on twitter).
It covers end-to-end, including ClickHouse as a vector database.
Especially seeing these days you can run a vector store on-disk if you have less than 10 million records, pull any free embedding model straight from HuggingFace and run on consumer hardware (your laptop).
Edit: another thought, skip lambda entirely and run the embedding job on the server as a background process, and use an on-disk vector store (lancedb)
In principle you could totally run this on a single bare-metal node, but most will not be doing that in practice.
why is storing the file as a FAISS/LanceDB on-disk vector store not "cloud native"? I am running this setup in production across dozens of nodes, we migrated all of our infrastructure off Pinecone towards this solution and have seen 10x drop in latency, and the cost improvements have been dramatic (from paid, to totally free).
I have a bit of an axe to grind in the vector DB space, it feels like the industry has gaslit developers over the last year or so into thinking SAAS is necessary for vector retrieval, when low latency on-disk KNN across vectors is a solved problem.
All that is to say, maybe there's a lot of money in the SAAS/big cloud space, and customers willing to run their own setup that requires tuning might not be willing to hand them large sums of money? Just theorizing here!
Oh also "cloud native" is like a marketing term vaguely saying "you can hook this into other cloud stuff" and it works with K8s/whatever cloud thingy.
Also I'm experimenting in further integrating things to reduce latency and most likely will publish another article within the month. Stay tuned.
Finally I somewhat agree that many of the players in the vector DB space try to push their cloud offerings. Which is fine, how else should they make money? And if latency matters that much to you, Qdrant offers custom deployments, too. I believe running Qdrant locally will handily beat your LanceDB solution perf-wise unless you're talking about less than 100k entries. We have both docker containers and release binaries for all major OSes, why not give it a try?
It also marks my first foray into using cloud services for a project. I've long been a cloud sceptic, and doing this confirmed some of my suspicions regarding complexity (mostly the administrative part around roles and URLs), while the coding part itself was a blast.
Not that I am against paying for a service, but the idea of writing my app against with a specific library against a specific platform makes me uneasy.
They have a github project but I think that is just the CLI + rust libs?
For a personal project it's just a bit much in my experience, especially since most personal projects can easily be served by a t3.micro.
UNIX is close to turning 50, and people are fundamentally paying as well as getting paid to make a written program loop to the beginning, instead of exiting. I think this is kind of wrong.
Many other commenters replying to https://news.ycombinator.com/item?id=36693471 are interpreting "complex" as "hard for me to set up." I think that's neither here nor there -- no matter what's underneath, you can always rig something to deploy it with the press of a button. The question is: how many layers of stuff did you just deploy? How big of a can of worms did you just dump on future maintainers?
It seems complex at first, but it is a lot more maintable and portable than creating aws infrastructure manually in the console. Once you leave your service to run for 6 months you will forget where stuff is, then in the worst possible moment if it goes down and you need to make some change you'll be franticly looking for aws docs... "can I create a synthetic canary and use the lambda I already have, or do I have to delete it and create it from Cloud Watch interface?" These kind of questions are the bane of Aws ops experience... And once you learn everything they "bring a new console experience"... So I prefer to learn terraform once and that's it.
Why terraform and not python with boto, cdk, cloudformation or ansible or something else? Because terraform is easy to port between providers (sort of), people who are not that good in python find terraform easier so you don't need "senior" people to maintain your code, finally it's a pretty "opinionated" about how w stuff should be done, so it's unlikely you'll open your project in a year and think "why in the world I did that!?", because all your tf projects will be very similar most likely. Also tf is mainly for infrastructure as code, there is no configuration management like in ansible... It is for one thing and it does it relatively well. (I have no relation to TF beyond being a happy user).
Cohere embed API https://docs.cohere.com/reference/embed
and AWS Lambda
Anyone knows what was used to generate the docs?
One of the engineers at work showed this to me a few weeks ago and ive been thinking of porting my Swagger docs over to it. I made a mental note to see if a free self-hosted open source version was available but have not gone back to check on that. Please let me know if you find suitable alternatives.
Edit: This is open source. Looks like the paid offerings are mostly to cover CDN costs.
EDIT: I looked at the source to get to this link.
[1] https://twitter.com/dylfreed/status/1651572024488218627
[2] https://xenova.github.io/transformers.js
[3] https://github.com/freedmand/semantra *(author of this tool)
[4] https://huggingface.co/sentence-transformers/all-MiniLM-L12-...