Storing a dataset of this size and making it available online is not inexpensive. Amazon has generously donated their services to handle both of these tasks; it would be foolish to turn them down.
EDIT: the state of the crawls are summarized at https://commoncrawl.github.io/cc-crawl-statistics/.
https://commoncrawl.github.io/cc-crawl-statistics/plots/craw...
Amazon makes plenty of money from people using AWS to processes the data.
[1] https://blog.sia.tech/announcing-skynet-premium-plans-faster...
Of course Sia's Skynet are package deals right now and I guess they're currently bootstrapping the network with users. Filecoin has no operational storage yet. Storj quotes 10$/Terabyte/month [1] so that will come out expensive.
1. https://www.storj.io/blog/2019/11/announcing-pioneer-2-and-t...
In fact someone could easily setup some IPFS nodes that fetch the data from the current host if requested over IPFS.[1] This way people could access it via IPFS and provide an alternate mirror of the data.
[1] https://github.com/ipfs/go-ipfs/blob/master/docs/experimenta...
The main benefits here would be
- Even if the source is unavailable there may be other copies on IPFS which would be transparently used.
- There may be some performance benefits in rare cases.
- If you are accessing this on a bunch of machines your IPFS gateways would handle downloading the source once, then automatically using the local copy from inside your network.
The maindownside is that if Amazon is donating their resources why bother with IPFS?