A recommendation engine that works solely based on your own watch history would require a totally different approach using algorithms that are years/decades/forever away from existing.
A recommendation engine that works solely based on your own watch history would require a totally different approach using algorithms that are years/decades/forever away from existing.
It does make sense that the job is simpler with everyone's history and inputs, if you can keep PII out of it. Are there any good references on recommendation aggregation algorithms?
Here's a simple proposal that keeps everyone's stuff anonymous and allows a rec service to pay bills.
Distribute the pool of everyone's watch histories/prefs/etc using IPFS, which is a DHT with static names and some storage. User clients put the share data into IPFS and register into a public pool. The client remembers their own hash. Rec services (there can be multiple) pull everyone's sets periodically and aggregate/indexe them. Users can then retrieve a batch of new recs by sending a small fee to the service of their choice along with their hash and the server will place some new recs into their bucket. If the recs are not suitable, they can try another service.
Even without this, I could just check videos you shared on Twitter, follow the same process, and have a decent chance of identifying you.
Anonymization is very very hard. Maybe federated learning is more promising here, keeping all your data privately stored on your own devices.
I think you're the one underestimating modern recommender systems, especially the ones built by advanced teams in big tech companies.
They absolutely use content features. Among other advantages, this counters the cold start problem (how will you learn about a new video if it doesn't get recommended to anyone).