This concerns me slightly, backup storage is a whole different world to real time data storage. Backups are write once, read occasionally, some people use S3 as a make-shift CDN, so constantly reading data.
Parity based repplication is great for backups, but would it not have performance implications if every request is reading from multiple disks/servers/nodes? I'm not an expert on hardware, but I would have thought being able to read an entire file of one disk is faster than having to put together pieces of data from multiple disks, anyone want to correct/inform me?
If you can offer me a serious alternative to S3 at a cheaper price, and open source software, I can't wait to try it out. I might sound negative but I just wanted to put across my first thoughts on having a look around the site.