A lot of commenters here on HN have already pointed that out already too. Unless they build a literal brick wall (paywall) around the site, that data can and will get scraped if the intention is to use for a model.
You could get it down to a science where you only scrape any new data whenever you train the next model.
[0]: https://old.reddit.com/r/reddit/comments/145bram/addressing_...