DigitalOcean lost our server
murze.be
murze.be
<begin quote> --------------------------------------------------------
We were loyal customers of DigitalOcean for over 2 years. We showed up to work one morning and had a client email stating that their website was down. We checked, sure enough. We tried logging in to our DI account and it said it was suspended. After searching around, we realized our CC had expired. No biggie, we'll just update it, turn the server back on and be on our way.
Nope. We had to contact customer service to get back in to our account. After updating the card info we realize that all our droplets are gone. We reach out to customer service again in which case they let us know that when a CC expires, an automated process kicks off and deletes the droplets. We've been a customer for 2 years, surely they could pick up the phone and call us. We've spent thousands of dollars with them.
Plan B, let's restore them from the backups we've been paying for. Nope! When they delete your droplets, they also delete your backups.
This is where we ask to speak to them on the phone and are denied. I then ask if they are insane and why they would delete someone's servers and backups via an automated process without a human at least checking to see if it is a loyal multi year customer who has a simple lapse with their CC expiry.
We had to find a backup on a developers machine from over a month prior and rebuild data using various megtods that took close to a week. In the end, DI gave us a $500 credit.
You get what you pay for. Anyone who uses DI for production or anything more than a hobby app is playing with fire. They do not care, they are apathetic, they will screw you over and the throw you a credit for their terrible service as a half assed apology.
What if I was on vacation, and didn't see this e-mail for a few days? shudder
In what world is automating the deletion of customer data a good idea?
The data destruction world :)
I'm not saying that excuses things, but it's far from "at the first sign of trouble".
My CC expired but I was on vacation and didn't have access to email. When I was back, the droplet and all backups and snapshots were completely removed. For the record, I was only 15 days late. Very sad. I tried my best to get the backups from the support engineer but no luck.
On one hand horrible situation to be in, at least could have called/mailed/SMSes/notified the customer before nuking the account.
On the other hand, it is a relief to know your data is secure and gone one you stop paying, no way they have a copy or keep it in archive/deep backup.
Having a backup and letting you to access it are two entirely different issues, and most of the time customer service staff don't know how the data management is being done by the engineering/operations team.
If DO's CC can't handle data management problems they probably will connect you to their engg/ops team. So I would assume they would prefer having it resolved ( by giving the data if they had ) instead of keeping it and let user rant/post about it( which they know will probably happen ).
The second paragraph hinges on the rationale that if customer data were available, then CS (Customer Service) agents would have made every effort to contact the ops/engineering team to retrieve those to avoid a social disaster.
Unfortunately, this isn't true in the companies that I worked for, which are big customer-facing MNCs.
First of all, for US consumers, the CS teams most likely are contractors in Philippines, Malaysia, or India, they are trained with scripted responses. For any issue that is beyond the scripts, their standard response will be "Sorry we can't help you because we can't do x", one reason is they have no incentive to escalate the issue and get oneself more work to follow up, other times, it is the management explicitly decided this is to prevent too many cases hit the ops and eventually, the engineering teams. It is not that the management is stupid or nasty, just that the 20:80 rule dictates 80% of the issues usually are raised by 20% of particularly picky customers. For the 20% customers who generate 80% of the revenue, each of them will be assigned an AM (Account Manager) who can get things done much faster and more efficient than the CS team. This group of high-value customers are often given access to VPs to make sure all their issues are heard and handled properly
Second reason is most teams in a big company will not concern anything that doesn't hit them directly. For example the CS team in my ex-company upon getting repeated complaint on an erratic component that didn't get bug fix for months, eventually the CS told the upset customer to fuck off and complain to CEO directly since it was engineering department's fault.
Unbelievable? Maybe, and that company remains one of the largest high-tech companies in the world.
I'm on Platform Support and while I can't access the actual hypervisor hardware like CloudOps can, we do not have "scripts", and we always do our best to help customers in any way we can. We also have a myriad of tools that we can use to monitor our platform and help troubleshoot any issues that may arise in our tickets.
My roommate and bestie is on CloudOps, and she sits right next to me at our HQ in NYC. While we have a great remote culture, all of us on support are in the US, with the exception of one guy in London (who started last week and is AWESOME).
Not only that, but I work very closely with all of the other departments here at DO, and our executives are extremely accessible as well. If we need to pass on feedback, we always do so.
We're also encouraged to go out of scope and do whatever possible to help out our customers, and I wouldn't change a thing. I love what I do, I love the people I help, and I love my coworkers.
* OVH.com — huge, great price point and performance
* Linode.com — also huge, favorite of many developers
* Vultr.com — owned by Choopa (bigger player)
* RamNode.com — doesn't scream "super modern", but reviews are great
There are many others, but those are off the top of my head that correlate with DO.
- 1andone
Vote against linode here, yes, they're large but they are also not all that good.
I don't understand that reasoning from customers to be honest. That right there, throws away the $5 and $10 plans. Having some one trying to contact a customer, just because their credit card expired, cost more than the actual hosting.
Top tier service does comes from companies with $5 a month plans. I'm sorry it just doesn't.
That's a decent amount of money right? It sounds like they were not using the $5 plans.
I would guess that if you want a hosting company to be proactive you need $100+ per month.
The VM hosting market isn't looking good at the moment, seems like each provider has some dark history at this point, it's really frustrating.
I'm quite serious. Our production environment operates on two cards with different issue & expiration dates from two different merchants and uses 2 different Cloud providers.
<begin quote> Hey Matthew,
I'm Zach, Director of Support here at DigitalOcean. Thanks for raising this topic.
I was able to locate your account, and I see that I was the one who granted the credit and followed up via email. I hadn't heard back from you until now, and I'm happy that we have an open line of communication. I'm hopeful that this thread starts a conversation, that it clarifies what steps we take, and we might even uncover a different solution that works well for everyone.
To start with, I want to be really clear, this is the absolute worst case scenario. I absolutely don't want it to happen to any customer, let alone someone who has been with us for such a long time. There's really a delicate balance that we must strike between customers who forget to pay and customers who do not want to pay.
Currently, our notifications for overdue balances are sent via email. If a customer account is unpaid on the first of the month, we send ~15 emails total notifying the customer of the situation, and subsequently power off servers 21 days after the account is on hold as a further way to gain your attention. At this time, droplets are removed from your account 14 days following power off, which is an increased amount of time from what you experienced. It was increased from 3 days based on past user feedback.
Why do we do it this way? In the past, we had no hold or suspension process for non-payment, which enabled bad actors to run for months and months without paying. From a business perspective, we made a decision to put a scalable process in place that limited how long an account could go unpaid.
As a support team and business, we are always willing to work with our customers who are unable to pay. We are always available via ticket, our contact form, and have made wide-ranging attempts to help users who aren't able to pay due to banking regulations: https://insidedigitalocean.com...
I'd like some input, so I'll publicly ask for feedback on questions that I've asked privately before. I'd also like any other feedback that we can consider on how to make this a more positive experience for everyone
-Knowing that we do not do phone support, what's the best way to notify you of a past-due balance? -Is SMS effective at times like this?
I would love to hear your thoughts.
Thank you, Zach zach@digitalocean.com
This updated policy seems a lot better; it would probably be great to ask people for secondary contact information such as SMS for emergencies.
At one of my positions we had stupidly put too many eggs in the basket of a single physical machine. Its disk controller failed in a way that it trashed the data volumes. I was unable to convince anyone that "move to Amazon" was not a one-step solution to "how do we make sure this never happens again".
To be fair the point of the cloud is that they deal with redundancy, HA, distributing to multiple datacenters, etc. through their services - but you need to use and understand implications of said services to leverage that.
But that doesn't stop the misconception.
The implicit (and explicit) volatility of cloud hosting should change that expectation with such services.
"Non-issue" is an exaggeration because it was potentially a catastrophe. But this is sort of the way these services "work." Servers/droplets/instances are ephemeral and replaceable, and their underlying data is not guaranteed in a failure.
These facts combined certainly make it easy for the younger generation to forget that redundancy and backup still don't replace each other (never did).
(Could be worse. Their first lessons could be like mine, in C! If anything will make you paranoid, writing C will.)
DO should be designing for failure: VM storage should be on a SAN. A single physical server failing should not cause data loss. This is basic stuff.
Does any cloud provider have SAN backed VMs? If any do, at what price?
SAN reliability ratings start at 99.999%.
I actually prefer the lower-level abstraction: if you want a lower failure rate (or higher speed), you can RAID together attached EBS volumes yourself on the client side and work with the resultant logical volume.
For ephemeral business-tier nodes, EBS gives you a few advantages, but none of them are that astounding:
• the ability to "scale hot" by "pausing" (i.e. powering off) the instances you aren't using rather than terminating them, then un-pausing them when you need them again;
• the ability for EC2 to move your instances between VM hosts when Xen maintenance needs to be done, rather than forcibly terminating them. (Which only really matters if you've got circuit-switched connections without auto-reconnect—the same kind of systems where you'd be forced into doing e.g. Erlang hot-upgrades.)
• the ability to RAID0 EBS volumes together to get more IOPS, unlike instance storage. (But that isn't an inherent property of EBS being network-attached; it's just a property of EBS providing bus bandwidth that scales with the number of volumes attached, where the instance storage is just regular logical volumes that all probably sit on the same local VM host disk. A different host could get the same effect by allocating users isolated local physical disks per instance, such that attaching two volumes gives you two real PVs to RAID.)
• the ability to quickly attach and detach volumes containing large datasets, allowing you to zero-copy "pass" a data set between instances. Anything that can be done with Docker "data volumes" can be done with EBS volumes too. You can create a processing pipeline where each stage is represented as a pre-made AMI, where each VM is spawned in turn with the same "working state" EBS volume attached; modifies it; and then terminates. Alternately, you can have an EC2 instance that attaches, modifies, and detaches a thousand EBS volumes in turn. (I think this is how Amazon expected people would use AWS originally—the AMI+EBS abstractions, as designed, are extremely amenable to being used in the way most people use Docker images and data-volumes. The "AMI marketplace" makes perfect sense when you imagine Docker images in place of AMIs, too. Amazon just didn't consider that the cost for running complete OS VMs, and storing complete OS boot volumes, might be too high to facilitate that approach very well. Unikernels might bring this back, though.)
Seems like they sell 50 GB of real SAN backed storage for just 3.6€ per month.
How can they afford to be so cheap?
Disclaimer: happy customer
Except that is not SAN backed storage.
From the website:
> We only use RAID-60 drive arrays with a minimum of 16 drives
That's somewhat scary. RAID-6 might be almost ok, but striped? No thank you. I bet they don't also have block level checksums.
Agreed.
> the customer should read them carefully.
Sure, but they shouldn't need to, as Digital Ocean should set expectations clearly.
Keep in mind Digital Ocean's <title>: Simple Cloud Infrastructure. Being surprised because there's an unsafe default hidden in a document somewhere didn't work out for MongoDB and it won't work out for DO.
They bought a hugely expensive SAN solution (HP I think). One of my questions was: What happens when the SAN fails? Well you see, because it has redundancy, that can't happen. Clever uh?.
First time it failed, everything gridded to a halt for two days. The second time they did better and had it running within eight hours.
That's completely different from suggesting 'SANs are magical boxes incapable of losing data'. Have you considered apologising?
But if DO had a multi-SAN setup with synchronous replication this thread wouldn't exist because DO's business model of "super cheap VPS w/ fast SSD storage" would have failed due to costs. Everybody wants Five Nines until they have to pay for it...
The lack of explanation is what worries me most - it leads me to think this might have been a case of "we forgot to replace a bad drive, then the second in the pair failed".
Fire happens, floods happen, electrical faults happen, mistakes happen. I can't really expect any human or humans to be perfect.
And many years ago it suffered a truck crash ( the truck crashed on the lamp post right outside the server company building and destroyed the telecom-related wires)
Forewarned is fore armed!
Keep it professional.
Though their should be another avenue for such things.
And when in doubt, if you're refund system is not capable of discriminating between the two, it seems to me wiser to err on the side of a less jubilant message.
For database servers you need to have procedures in place to quickly switch the production application over to a database you just spun up and migrated the data from a recent backup to in order to test your data backups. You want this to happen as smoothly as possible in the case of a failure. Keep data backups in three different places and test your procedure on all of them.
On a production system, do not keep the database on the same server as the application.
If the server goes away, then even once I redeploy, I've lost all those cron jobs. Who knows what will happen then.
On systems I build, cron jobs are added as part of the deployment process and managed as code, in the git repository.
That's just one example. Others include iptables rules, FTP server configuration, startup scripts, and such. If it's not in your codebase, it's unmanaged state on the server that you will lose if your cloud provider takes a shit. All of that needs to be managed as code or as configuration and deploy scripts should contain idempotent commands that add them if they are not present. I consider shelling into a server for any reason to be unideal and will think of ways to avoid it.
If you are asking about database being on the same server as the application, it's because it makes infrastructure management more difficult. Application servers are set up quickly, torn down quickly, and work as soon as you deploy. Database servers are fortresses, you don't build new ones and tear them down half as quickly, unless you've got scaling requirements.
It is tempting to want to put data on the same server as the application early on in order to save on hosting costs, with the understanding that you'll separate them later, but that's a rookie move. It's much less work to separate them now, when there's nothing riding on it, much less headache.
That said you could easily use the same advice for that data. But in some cases the application itself needs to be on the same server as the database. Think if you were building postgres for instance...
If you can't afford to lose it, then it needs to be appropriately managed so it can be recreated in the case that it's lost. If you can afford to lose it, then it's not really state and you can forget about it.
I'm pretty sure the usual understanding of the terms "business tier" and "database tier" is that the "business tier" is a bunch of ephemeral nodes, and the "database tier" is where-ever the data from the "business tier" goes to be persisted.
Configuration capture is a big thing, and it separates the adults from the kids. I've had hardware failure of my dev box twice in my career, and when I tell certain people that I have to rebuild my machine you can see their faces blanche. These are the people I know I haven't converted yet.
Everything important, except for the handful of things I'm actively investigating, is always stored on shared machines. Version control. Wiki. Some sort of artifact repository. The instructions for getting an environment set up with version X of our software are also stored in one of those places, and have been vetted by every new developer and some of the QA team. I can have my machine back up and running in a couple of hours, and most of that is waiting for downloads, if we didn't have the presence to store copies locally.
Why do I care about this stuff? Seems sort of OCD on the face of it, and maybe you're right. But sooner or later you're going to have a high severity issue in production that should be an all-hands-on-deck affair, and if you haven't done this work, losing a hard drive on a dev laptop will be the least of your worries. Everyone busy doing work on X+1 needs to be able to get back to version X and all of its dependencies in under an hour, and by themselves, because the people who could help them are most likely on the front lines of fixing the bug as fast as possible.
What's more, someone probably needs to get versions X-1 and X-2 running to figure out if you need to warn people using older versions. So that's getting people running 3 or 4 versions of your software autonomously, so that you can identify repro steps, long and short term mitigation strategies, verify that they work, formulate a bulletin and provide patches for people. Not only do you need to get your configuration captured/documented, you need it captured 2-3 versions of your software ago, which means you need to start thinking about this stuff pretty early in your project.
If that takes 4 hours you've hardly spent any money at all, and you're running your cluster the way a lot of people already do.
I agree that DO handled this situation appropriately, but they should warn people a little more deliberately when they sign up and create a droplet. A checkbox along the lines of "I, user, understand that I'm responsible for backing up my server", with a few links to some how-to articles of the excellent quality that DO is known for, would suffice and empower without reducing the conversion rate.
I don't host anything I couldn't live without though, so I'm not really worried.
My new model: client wants a pretty responsive theme, I get the HTML version of it. Turn into a template I can use to spit out a static site (from my homebrewed CMS). I tell clients your monthly support charges are $0. If you ever need updates I charge hourly (most never call in months). Some clients want "a list of something" say products. If small and won't grow: a static csv file that also gets munched by the homebrewed CMS. Large list? A BAAS service with API.
I really don't care if the servers are gone/hacked/ransomware. Git clone, go.
The client (who doesn't want to contact the agency again)is given instructions how to update which makes them happy. And then such projects spread like wildfire.
/rant over.
For added protection, you can take regular snapshots. You only pay to store the diff from the last snapshot (so go ahead and snapshot often), and snapshot storage is geographically distributed.
(Note: I have no idea what EC2 does, maybe it's similar.)
Amazon EBS volumes are designed for an annual failure rate (AFR) of between 0.1% - 0.2%, where failure refers to a complete or partial loss of the volume, depending on the size and performance of the volume. This makes EBS volumes 20 times more reliable than typical commodity disk drives, which fail with an AFR of around 4%. EBS also supports a snapshot feature, which is a good way to take point-in-time backups of your data.
https://www.joyent.com/blog/on-cascading-failures-and-amazon...
https://www.joyent.com/blog/magical-block-store-when-abstrac...
https://www.joyent.com/blog/network-storage-in-the-cloud-del...
Besides Joyent, I think DigitalOcean and Vultr do it right, with RAID-backed local storage and a backup option.
I agree that if you have a large-scale sophisticated operation you will probably want to choose local disks and handle availability your own way (and GCE provides those options). But for small-scale operations that can't justify as much engineering, abstractions like persistent disks and automatic migration will save a ton of time and avoid data loss. Ironically, Digital Ocean seems to target this smaller scale, yet they've set the wrong defaults for their target market.
I think that's a poor lesson learned here. Were this me, I would have said:
"Luckily, all of our sites run on several servers, access data in a shared, replicated cluster, and a small shell script I wrote kept me from writing this entire blog post."
IaaS has only surfaced what has always been true: your data lives on little physical things that are screwed into a thing and goes through a controller that could fuck up due to cosmic rays.
Do better for your customers.
Backup servers no matter where hosted, because two is one and one is none.
1. Never have a single point of failure. Relying on DO for Server+Backups is putting your eggs on one basket.
2. Your server state should be programmable. This is not quite easy for complex configurations. But today, DO has an API, we have Docker, and quite modern deployment tools.
Here is setup:
1. Github for the server state. Basically, a repository to configure and deploy my infrastructure.
2. Enable DO backups in case I mess up something and want a quick come-back.
3. File backups through Tarsnap. Since I use Docker, I have a volume container. Backup the volume container with Tarsnap.
Pretty much any use case where I'd use something like DO being unable to create new VMs, etc. is the same as an outage.
I asked why as I work for DO's operations team and I wanted to know what your concerns were. Do you have a specific instance of a problem? What you're describing would be considered a MAJOR outage for us and we have not had one in quite some time.
* Jan 6th
* Dec 11th
* Dec 4th
* Nov 27th
* Oct 10th
* Oct 2nd
That is annoying enough I'd want to be able to failover if it exceeded 15 minutes. I get it might not be an issue for anyone else but failing over is less complex than mitigating latency issues with droplet creation or whatever.
I get "regularly" to me might mean something different to you but if the 2 DO DCs I was using are hit literally every month with a problem of some kind...it seems regular to me.
I award OP the "Most Accurate Title Of The Day" Award.
None of these would be an issue if you've scripted server recreation from scratch and permanent data is stored in high availability shared databases and services like S3. If you're relying on weekly backups to save your server state and customer data you're just asking for trouble. I like that Heroku recreates your server once a day to force you to do this properly.
> Three backup slots are executed and rotated automatically: a daily backup, a 2-7 day old backup, and an 8-14 day old backup. A fourth backup slot is available for on-demand backups.
Obviously, you want some offsite backups rather than completely relying on DO but that should be for disaster recovery (massive failures in DO's DCs, etc) rather than a more routine one.
Because all those earlier breaches(which you are referring to) revealed customer info not their data. If this was your point of user info, I wouldn't have asked.
Having something 7 days old might be useful in some cases (especially e.g. DB dump) but it is far from 'good enough'.
I am not sure what's the reason why they do not provide more recent backups (for additional fee, obviously).
Would it be so complicated that the engineering effort required wouldn't make it worth it? Or such a service wouldn't be popular enough among DO users?
After learning that Digital Ocean does no backups of their own (even for critical hardware failure on their own side) I've enabled weekly backups for an additional $1 per month.
Glad you posted this so that I know that option is available and really necessary. And who wouldn't pay $1?
Edit: It costs 20% but I only pay $5 for my personal server.
I just don't even know where to begin. Is this something I could stand up locally and push to like s3, spin up a new server and install 1 thing and have it pull the configs and software installs down?
Any suggestions?
My opinion is ansible. Basically python + yml, you list out steps to set up your server. Takes a few minutes longer round 1, but you can remake the server in 60 seconds in 6 months. Lots of ansible tutorials around, try it out..
user data lives in redundant databases / storage service
server setup lives in a repo via bash scripts or ansible etc (and bake that into an ami or container)
setup your launch-configuration for your servers to setup all of the above and you can destroy and spin up servers on a whim
Yes. You could use shell scripts, Ansible, Salt, Puppet, whatever. I'm fond of Ansible or shell scripts, depending on complexity.
Start here: http://www.opsschool.org/en/latest/config_management.html
A better solution would be wonderful.
Also backblaze's b2 is pretty damn good/cheap. I use it as a secondary backup.
Even damn Netflix preserves data for 10 months. Deleting backups is making 100% sure that the customer doesn't come back to you. It's amazingly stupid.
Unless you're as smart and as big as Google, use the damn phone. Don't use the word automated at all. It's not automated if one fool made a decision and another one wrote the script.
That said, I'm a fan of moving away from RAID. It's just more complexity in the face of a movement towards simpler architectures. The one benefit of a good RAID card: stupid fast, safe writes with a BBWC.
I can re-implement the infrastructure of my VMs fairly easily (thank you Ansible), but backing up content outside of the provider's built-in options is something I haven't played with yet, and would obviously be the crucial piece.
Note that we only back up specific things like databases, logs (central syslog server) and data directories. We use Puppet to configure our boxes, and consider everything to be expendable that isn't data.
And yes, I do a nightly backup of databases, webroot, /etc and some other directories.
So, all this strikes me as "lol you didn't know this?"
EMC and Netapp are of course another matter, they're great. If you can afford them.
I'm a fan of ZFS, but this is the first time I've heard this claim. Do you have any links to articles/stories talking about it?
Server seems down :-/
adults acting like children. nothing more, nothing less.
take responsibility for your own data. back it up yourself.
I run Redundant disk controllers and hard drives on ZFS. And it's cheaper than a decent RAID system.
First thing I do with a new server is disabling RAID.
I hope you see what I'm getting at. While I love ZFS and think self-hosting can make a lot of economic sense, without knowing all the variables you can't really make a fair comparison.