a) Massive data volumes (~100 Gb - 1 Pb/project)
ai) This means that data is typically stored on limited access machines like HPC clusters
bi) This also means that shipping this data around is financially expensive, and cannot be supported purely by small client machines
b) A low number of seeders; scientific data is not exactly popular, and there may be network restrictions on uploads through the typically used networks;c) The requirement for a data legacy; torrents are fantastic for ephemeral data (e.g. operating system builds), but are terrible for data that must be archived and kept for potentially decades to centuries.
It would be challenging to find a solution robust enough for CERN type data but also simple enough for an n=3 undergraduate research project (that may have yielded some interesting results).
I don't know what the solution is there. My intuition is that university libraries could be involved, and that a data librarian could help you get your small study into shape or be embedded at a percentage effort on a large study.
It only adds. I don't understand how it subtracts.
Hoping IPFS makes it someday because the idea is great.
> IPFS or torrent are the best options for distributing data
And suggesting that IPFS is not a good option for distributing scientific data due to its complexity.
In this discussion it means "distributed" in the "made available" sense.
The parent did not use IPFS because it didn't work for them. So no, the data was not distributed via IPFS.