There are a bunch of possible areas to circle or ignore when making an ML-capable database of some sort. In rough order of data complexity:
1. Embeddings (context-free vectors, just an ID and the vector)
2. Metadata + Embedding (source data, JSON)
3. Binary Data + Metadata + Embedding (add documents)
Then there are tooling questions: in this matrix you'd want to decide if you're going to allow inference, and if so, will it be arbitrary, service-based, etc. against the documents, and if so, how will you store the results?
I'm curious how you're thinking about the design space. The embedding-only route is conceptually appealing because it's simple. In a larger engineering project, there's a tension between "where do I keep all this data," "how do I process and reprocess all this data", and "where do I keep the results of all the processing", and to me there aren't clear bright-line architectures that seem "best of".
Put another way, 15 years ago, we went memcached -> redis 1 -> redis (whatever it is now), and at the same time, we went mysql/postgres/oracle -> nosql json stores; today all of these have relatively well-defined use cases, (and for most of them sqlite is the best choice, obviously).
How are you seeing the ML db scene playing out, and where do you think the sqlite of this space will land on architecture?