I tried running this on a H100 and got 190ms compared to Jev's 170ms. Maybe I set it up wrong?
4,299 karma · joined April 23, 2018
Hoping to implement a simple RL loop here and optimize whats generated by the LLM to create the perfect slop machine :)
for clarification :)
In the meanwhile I tried at least 100+ variations of trying to train this model to be SOTA most of which led me down getting better data which I believe I do, just needs more tuning still.
so tldr most of the issue the author has is against the person who made the library is the design not the implementation?
I think milvus, quickwit, and pinecone are geared more towards enterprise and are hard to use.