838 karma · joined June 14, 2014
What are you even talking about, "dude"? I don't think you have the background or knowledge you seem to think you have to argue about this space. It's OK, as you blithely pointed out in your earlier post, there is a place called medium where you'll find likeminded folks that will eat up your drivel.
E.g., see the nice work on the PAM library (https://arxiv.org/abs/1612.05665). Ideas from this work were used to build lots of cool things (immutable graph data structures, segment trees, databases) that are very fast, and all immutable.
The result looks very interesting, and the blog post is well written (e.g., I did not know about the prior work re. Grover's algorithm and pendulum systems).
The blog post is also based on a recent FOCS paper, and the authors are reputable people in CS theory, if that convinces anyone to take a closer look.
The other useful resource I will mention is "soulseek" (https://www.slsknet.org/news/) which is basically old-school direct downloads, and includes a lot of obscure music that you won't find on AppleMusic/Spotify.
Happy trails!
Users also make wonderful charts on this website. also, hello from a fellow what refugee!
They're not OpenAI embeddings, but they are realistic, and much larger (number of vectors).
I think many production embeddings at non-OpenAI companies will use lower-dimensional vectors than 1536, so it makes sense to focus on non-OpenAI embeddings as well in your benchmarking.
I’m also curious if there is a way to not store everything in memory for pgvector. Is that possible?
Lastly, what is the parallelism story? Is it just using a thread pool under the hood? OpenMP?
Understanding if pgvector plans to support point insertions and deletions is also important in practice.
Come on man. Having seen the inside of big-tech-TM and the senior engineers there, yes they are fast and good, but they are not 100x better than the new guy. Maybe 3--5x at best.
Anyway how do you train good senior engineers? They don't just pop up out of thin air.
But really we shouldn't be talking about math errors and falsification in the same category.
E.g., in math / theory, if someone has a history of making big blunders that invalidate their results, you will be very hesitant to accept results from a new paper they put on arXiv until your community has vetted the result.
So yes, I do trumpet science as a model of a rational, self-correcting social enterprise, at least in CS.
Other sciences like biology and psychology have some way to go.
FWIW, most top-ranked CS conferences have an artifact evaluation track, and it doesn't look good if you submit an experimental paper and don't go through the artifact evaluation process. Things are certainly changing in CS, at least on the experimental side.
It's also possible that theorems are incorrect, but subsequent work that figures this out will comment on it and fix it.
The scientific record is self-correcting, and fraud / bullshit does get caught out.
You can google for more results / papers.