I also curate a list of blogs as part of my blog discovery/search engine/reader project: https://minifeed.net/blogs
I also curate a list of blogs as part of my blog discovery/search engine/reader project: https://minifeed.net/blogs
For OPML, tracking the enhancement here: https://github.com/blogscroll/blogscroll/issues/198
The related posts are generated with the help of a vector database (Cloudflare Vectorize [1]). I take post titles and descriptions as input, vectorize and store the results in the vector DB. Then for each post, I find 10-15 related posts and store the results in a separate table (it's a plain one-to-many SQL table). This is done because querying the vector DB is not super fast, and I wanted Minifeed to load under 500ms (according to my status page tracking, users in Europe for some reason experience the fastest loading times of 100-200ms! [2]). I also set up a scheduled job which regularly updates the relations (since there are new posts every day, and some of them may become "related" to existing posts).
I've been meaning to write a blog post about the overall architecture of Minifeed, there are lots of small components related to RSS parsing, updates, caching, etc.
1. https://developers.cloudflare.com/vectorize/ 2. https://status.minifeed.net/
Surprised to see my blog[1] there :)