Retro on Viberary
vickiboykis.com
vickiboykis.com
It’s interesting to compare notes and implementation details. I have many of the same notes and faced very similar challenges in my own work haha.
The most interesting challenge I found was touched on in this post. Semantic searches presented with the typical text box input generally mislead users into treating the search like a Google search.
Searches like “beautiful video game” and “sci-fi thriller” technically work, but they don’t perform as well as just describing a beautiful video game or describing a plot overview of a sci-fi thriller.
Semantic search kind of requires a different kind of search query that many users don’t quite understand yet.
Maybe there’s an opportunity for an intermediary model to “translate” queries, like how recent image generation models take your image prompt and generate a text prompt that’s used for the image generation instead.
It works if there actually is a game out there you want similar games to, but the text input is useful for finding something when you don’t have an example offhand.
Alternatively, you could mean pool (average) the document embeddings for the most popular click results for a given query and use that as the query embedding.
You can see the project here: https://azstatic.danieltperry.me/steamvibes/build
I’m not 100% satisfied with the performance of the search, but it is what it is. I wanted to get what I had out instead of perfecting it all, since there's a ton of different things I can do that all have ambiguous amounts of impact to the results.
I might switch to using a different embedding model in the future since the current one I'm using seems to be fairly dated by this point.
For embeddings, we chose BGE, but found they seemed to overfit on Beir. CohereV3 and Voyage appeared to perform better in practice.
For retrieval, we used OpenSearch + ZillizCloud Serverless + Cohere ranking, providing us with maximum flexibility and search effectiveness.
To be mentioned, build a evaluation dataset is important. This helps me to improve the quality.
they actually have custom prompts for each dataset being tested.
Question would be, if you haven't seen the task before, what is a good prompt to prepend for your task?
IMO e5-mistral is overfit to MTEB
see the retrieval tab
I would assume Amazon's product suggestion would be SotA for reccomending a book based off of another. It is a recommondation system and while it uses semantic search there are many more ranking signals it uses.
You're talking about the same Amazon that will blindly give me recommendations for other products that fill exactly the same niche as the one I just bought.
"Oh, you just bought a generator? You probably need a second one too, right?"
I assume this might be behind the state of the art, then. As of a year ago OpenAI managed to get an embeddings model that can do both without any special flag or dual model to treat queries and documents differently.
For example, on this project, looking at the generation of training data [1], it seems like what's actually being generated are embeddings on a string concatenated from each review, title, description, etc. [2]. With the max_seq_length set to 200, wouldn't lengthy book reviews result in the book description text never being encoded? Wouldn't this result in queries not matching against potentially similar descriptions if the reviews are topically dissimilar (e.g., discussing author's style, book's flow, etc. instead of plot).
[1] https://github.com/veekaybee/viberary/blob/main/src/model/ge... [2] https://github.com/veekaybee/viberary/blob/main/src/model/ge...