46 karma · joined May 11, 2024
Maybe at some point I (or someone) can create a server version repo so people have the option to choose.
I'm pretty new to embedding so my understanding may be a off.
You can also specify more granular in human words what you're looking for which is a big bonus for me personally.
I just got tired of manually inputting the data and wanted a more automated approach. This recommendations system isn't extracting loads of data yet (like how often things are watched etc..) but instead a more birds eye view of your library and watch history to analyze.
If a model was trained 6 months ago for example it will likely have some info on shows that came out this month due to various data points talking about that show as "upcoming" but not released. Due to that it may still recommend "new" shows that have just released.
All that being said, I have to imagine that suggesting shows that have just now been released is likely the weak point of the system for sure.
The general idea is I generate a prompt to feed into the LLM with specific instructions to use the Sonarr (for example) library info which is passed in as a list of series titles. I also provide some general guidelines on what it should take into account when choosing media to suggest.
After that it's in the hand of the LLM, it will internally analyze and recommend shows based on what it believe someone might enjoy based on their media library...Given that every LLM model is different, how they internally decide what shows to recommend will be unique to them. So if you have bad suggestions etc..It's best to try a different model.
it provides nice flexibility but in reality my control of the actual recommendations are limited to the initial prompt that is passed in.
At some point though like you say, it's going to become ineffective and you'd probably want to use the "Sampling" mode that is available to only send a random subset of your library to model as a general representation. Though how well this works on massive library remains to be seen.
I haven't looked at integrating Overseer yet but that is a good idea as well and worth a try at implementing. I'll be adding that to my list, thanks for suggestion!
As for the largest library, I only really know of my own which is around 250 series and 250 movies. Not small but not huge. Passing all of that info is fine enough, but I'm also curious how truly massive libraries or watch histories are handled.
I imagine you would hit the LLM token input limit first if you had thousands of series and movies. Definitely need some further testing in those cases.
Jellyfin is definitely on the list to be added, it's probably next in line actually. If it's as simple as Plex integration I'll be very happy.
Edit - Support Added