That being said this is a great addition! Looking forward to trying it out when it reaches my datacenter. And I'm also looking forward to see what their next big project will be.
That being said this is a great addition! Looking forward to trying it out when it reaches my datacenter. And I'm also looking forward to see what their next big project will be.
Most of the great new features proliferating on other providers seemed designed to encourage vendor lock-in. On my current project, I made one concession to my usual aversion to proprietary lock-in and went with AWS S3 for storage of user uploads (also looked at Azure's offering but my BizSpark application was denied, so meh), since there really is no equivalent. I'd thought about using Linode's beta block-storage (which would have probably been more performant), but I was afraid I would be forever fighting with managing volumes and nfs mounts (or massively overbuying capacity).
In general, it's really hard to do at-scale network-backed storage - by the time your applications get access to the file system, there are a myriad of abstractions that aren't always receiptive to the idea of the network "going away", or even a modicum of lag. On top of that, in order for it to be profitable, you need to work at a massively-shared scale. This means expensive SSDs, servers and switches that require a lot of capex and no guaranteed revenue because it's a new product. For us, this meant building entirely new network architecture in some places to allow the massive amount of data being shared in the storage cluster and across VMs to not overwhelm existing traffic, etc.
In order to create the reliability in network and persistency that your normal application desires, you need extremely strong consistency and low latency. Every replication strategy (replication and erasure coding) requires each write to touch more than one SSD/HDD/NVMe device in order to acknowledge the write, and that all needs to happen in a shared system with an immense amount of contention, every time.
It takes a while because you only get one opportunity to get all of this right - it's one thing if the network has a few more blips in a month, or if there's a bit more CPU contention than you'd like, but you absolutely can't lose peoples' data.
I can understand why companies are so hesitant to do this - there may be technical debt in their software/network stack that makes it very difficult, or they may not want to proceed unless they have the right set of experts working on the project.
We've had very good success with Ceph for block storage and a fairly rough time with it for object storage. We're currently doing our best to try to improve it (both our own upstream contributions and collaborating with RedHat).
From a technology standpoint, I think it is very interesting and for the most part has a lot of very good engineering. However, it is fairly complex and even today it's very easy to have a hard time with it when starting out. You really need to pay attention to every detail and your hardware selection is extremely important.
It is extremely resilient, and goes to great lengths to preserve your data. Ceph can be performant, however that requires very good hardware and network.
My experience is limited up to the Jewel release (we haven't upgraded to Luminous and we are not planning on using BlueStore anytime soon).
Sage's talk at LCA covered the work they're doing here; https://www.youtube.com/watch?v=GrStE7XSKFE
But yes, at small scale Gluster is still a lot easier to deploy and run.
There are options that will perform better, but they are almost always considerably more expensive than FOSS, and all have their own weird scaling quirks.
With the launch of BlueStore a few months ago as well as improvements in erasure coding, I wouldn't hesitate to take a look at it again if I was starting a new project.
In some sense, "cloud native" architecture has shifted all the hard problems into storage by treating all non-storage resources as transient. So storage is the one place where persistent state exists and you can't just reboot it into a clean state.
But yes it was a hard road to get there, particularly in the days when Linux's 10Gbps drivers and btrfs were less good than they are now, and we also needed to write a new NBD server for all the live migration to work smoothly (https://github.com/BytemarkHosting/flexnbd-c)
For instance, I use OVH object storage in conjunction with an image hosting site on Linode. The latency from OVH Canada to Newark is small enough its pretty seamless and if OVH Canada goes down, I can use an EU location with higher latency. Linode fails over to their London location (assuming there is enough availability with VMs, the spin up may be automated but I run 0 webservers there 99.9% of the time).
Personally, I wouldn't pick DO because they tend to have poorly disclosed problems with their "new features" that you really only get an answer to if you contact support or dig through their documentation. For instance, DO's object store doesn't handle index files at all but everyone else's does. So if you try to switch, you almost immediately end up with "eh...srsly?" moment.
Linode and other hosts at least deploy the standard feature set when they expect money from you.
It's... pretty frustrating, so I'm looking for a new provider with a simple API.