Myscaledb: Open-source SQL vector database to build AI apps using SQL
github.com
github.com
A database that scales to billions is cool. Feels like k8s though. I'm curious who needs it.
Right now I'm interested in in-process retrieval options. We're going to use our own document databases anyways, so having the retrieval database in it's own server just adds an extra layer of complexity.
Cool, but does it actually return the correct results for these SQL statements, especially when ORDER BY is concerned?
I.e. does it somehow have a way to get a recall of 100% from its indexes?
https://clickhouse.com/docs/en/optimize/sparse-primary-index...
Considering that embedding vectors represent a lossy compression of the original text or images, is achieving a 100% recall necessary? I am interested in understanding its practical implications.
Disclaimer: I am an employee at MyScale.
For the app, maybe not. But as a database absolutist, I think you must be able to dump all rows of a table with
WITH
limit_result AS (SELECT *, {similarity} AS metric FROM table ORDER BY {similarity} ASC LIMIT 10),
dist AS (SELECT MAX(metric) AS max_m FROM limit_result)
SELECT *, {similarity} AS metric FROM table, dist WHERE {similarity} > dist.max_m
UNION ALL
SELECT * FROM limit_result
... assuming that the ordered values are unique across the table and fully sortableA recall of <100% may skip some rows in the limit_result, which then also won't show up in the main table's scan result, thus potentially corrupting a data dump process that uses sorted output.