An index is a derivative-enough work that copyright doesn't apply.
An index is a derivative-enough work that copyright doesn't apply.
There are 195 other countries you could scrape from though.
As for whether an index constitutes copyright infringement, that depends on its construction, but in the US generally not.
The terms of service bind the legal entity that is accessing Reddit’s servers, whereas reproduction of the data is governed by copyright law. Microsoft would still have to abide by US copyright law but there’s some wiggle room in terms of who is doing the scraping.
The idea of ignoring robots.txt when it is abused for greed sounds okay, but it would be much better if the industry was in a spot where we respected each others' requests, and peers don't make requests that put legitimate, honest actors at a disadvantage compared to rouge crawlers