> Any contents blocked by the current robots.txt is removed retroactively from the entire 2013-2024 range of the training dataset
Why not check historical versions of the robots.txt (e.g. archive.org) and contain the retroactive cutoff to a certain date range, parsing the robots.txt accordingly? That might increase the corpus size within legal and fair use boundaries.