I wish there were more details about how this list of domains was compiled.
This is the explanation on the Gitlab:
> We were testing an AI that could show some basic emotions about internet content, and turns out it was very precise at getting “annoyed” by ads and “unsolicited” third party connections…
> From that, I forked our own project and tweaked it in a specific way to basically only focus on ads, trackers, etc. and act like a web crawler, turns out to be very effective!