Future of Cloud Storage: An inside perspective from an IaaS vendor
cloudsigma.com
cloudsigma.com
While I don't claim to fully understand the advantages brought by block storage vs. file based storage, NAS storage should be usable for most applications. If this is true than the clustered or scale out solutions offered by IBM (SONAS), HP (IBRIX) and EMC (with the recent acquisition of Isilon) might be a suitable off the shelf solution matching your scenario (probably there are some open source solutions also that I am not aware of).
The way I understand these solutions you can basically scale capacity and performance independently which in turn translates itself in the flexibility of achieving the ideal balance between capacity and failure point distribution/redundancy.
What do you guys think of this setup? Would it achieve your goals?
I'm coming from a corporate environment where having an off the shelf solution that is supported by some major company is as important as the solution itself.
So, we believe that a move from our current set-up to distributed block storage will actually REDUCE costs overall no increase them. We will of course have the added convenience and elimination of single points of failure in storage that distributed block storage entails.
Best wishes,
Patrick CEO CloudSigma
Best wishes,
Patrick
When distributed storage systems like ceph have been around and in production for 5-10 years, I'll re-evaluate. but for now, they present too much risk of data loss in terms of bugs and admin error.
there's also gluster, which is much closer to my own standards in terms of 'time in production' but even so... local disk is simple, and when it fails it fails in a non-spectacular way.
Of course, the other problem with distributed filesystems is that it makes having a good network /much/ more important. Hell, right now I could get away with 100Mbps, so on a gigabit network, I can have some pretty serious network issues before anyone notices anything is wrong. For a widely distributed storage network? even at my current scale, I'd at least need 10G interconnects between the switches, and god help me if there was a network glitch.
As pointed out in the article, using distributed block storage means you need a pretty high performance storage network and low latency is just as important as high bandwidth.
We already have a physically separated storage network that runs in parallel to our public network. This is used primarily for drive traffic over iSCSI. We have a standard gigabit redundant network for this. As a physically separated network it means we don't get data integrity issues caused by DOS for example.
Moving to distributed block storage will mean we will upgrade our storage back-end to either 10Gbps Ethernet or Infiniband. The advantage of Infiniband is that as a cloud provider we will have essentially a large grid which is what Infiniband is meant for. We can also use 40Gbps per port so with dual networking can go up to 80Gbps relatively cost effectively. Just as important is the super low latency. Suffice to say we are ahead of the software on this one :-)
I think its important to point out that as storage moves to these sorts of systems it will be come increasing difficult to replicate such a setup in a redundant fashion on dedicated hardware. As a cloud provider we spread the cost over many customers.
Best wishes,
Patrick, CEO, CloudSigma
On the other hand, infiniband is fucking awesome. DMA over the network, anyone?
There are examples that have been in production for at least a dozen years in commercial operating system deployments.
The following is multi-host RAID-1 storage, known as host-based volume shadowing (HBVS) storage:
http://h71000.www7.hp.com/doc/84FINAL/ba554_90020/ba554_9002...
Getting DRDB working is certainly interesting and not a small project, though dealing with the many and various error cases is the truly entertaining part of the effort. Different devices and different failures can and eventually will toss back errors, and with different timing. And from experience with HBVS, these errors can and do shake out in production, and that's bad.
Patrick, CEO, CloudSigma