Show HN: I built a vector database API on Cloudflare
github.com
github.com
This also uses a fixed embedding which will not be compatible with all machine learning projects that people will want to use a vector database for. The chosen embedding only supports text so making an image search for example wouldn't be possible to do.
Most vector databases are using some local or external provider to get the embeddings and then using some storage engine to store and retreive the embeddings. Whether it's pgvector leaning on postgresql, or chromadb on sqlite, or pinecone originally being on rocksdb (I believe they've now built their own engine).
This is no different, and is still in it's infancy, so one presumes they might add support for other methods of getting the embeddings much as chromadb has.
retrieve != query
this project is extremely simplistic in regards to its vector search tech. pgvector is an open source implementation of an _index_ (multiple algos actually), this uses Cloudflare's completely proprietary index with a single call.
Athena is already a name in the cloudy database segment though - Amazon's managed Presto offering
Very cool project but unfortunately the name is immediately a trademark issue.
https://aws.amazon.com/trademark-guidelines/
I would expect a cease and desist at some point.
Fighting for ranking on Google. Conflicting tags on stackoverflow. Confused users sending you complaints about someone else's product. And so on.
The first commit to AthenaDB was on Dec 11, 2023.
Is the prior use to a trademark the important date or the actual date of registration?
Let's say you have a team of 5 data scientists/developers who are working on a collection of GenAI features/tooling. Does it make sense to have one single vectordb where all documentation is embedded and powers all the apps, or do you make a bunch of niche databases that are tailored to the service?
Also, one of the things i've noticed is that these databases seem less optimized for update operations so when user #1 embeds and saves 100 documents then user #2 does the same, with 10 overlapping - I'd guess that doubling of the similiarity space would exclude new documents. How are people handling that?