Digital Ocean needs to start showing outages on their status page
medium.com
medium.com
It would be unreasonable to have a company-wide status page that constantly lists "some customers are experiencing some problems". That's not the point of the status page - the status page, as the author suggested, is there to highlight issues that are affecting a significant section of the customer base.
The right thing for Digital Ocean to do in cases like this is to allow you, in your private dashboard, to see the problem and follow up on a master ticket for escalation and resolution.
While we've listed every outage affecting >1 customer since 2004 on our status page, the issue is always expressing each outage in a way that allows a customer to identify that _their_ server is affected by a particular entry.
That sometimes involves knowing that their server is in a particular data centre, or that it is connected to a particular switch, etc. and we do our best to make sure that people can identify their problem, if nothing else than by timing, i.e. making sure we list something ASAP.
But status pages are still very useful once people start calling in - support can positively identify that yes, you are affected by this problem and you can track progress at that URL. If we keep updating as we promise, bang, that's one support call (at most) per affected customer.
So even for a minor problem - broken switch, VM host machine, power failure in a rack, that's worth listing from my point of view.
As I said above I'm looking to tie our databases and some basic network monitoring into it this year. That way we can proactively notify people affected by a particular problem, as well as continuing to list even small problems publicly.
A lot of the other suggestions seem to centre around extending your status page to be more personal to the end user, this is (to some extent) the route Amazon AWS has taken in allowing you to see which instances are scheduled for retirement (because they lack live migration capability, a la GCE), etc.
I note that Amazon now sends out maintenance e-mails advising when certain IPsec connections will go down for them to perform upgrades, this is also great, and others should copy.
This leaves the 'globally visible' bit (i.e. http://status.aws.amazon.com/) for critical outages affecting a large proportion of your customers.
Is it wrong to provide a lower-cost service?
Some people run their email server on these budget IaaS. Some prefer to host outside of Google or Amazon's power. so where else should they host their own server? Home?
Ideally if we have continuous streaming backing up a node, then when the host machine failed a second machine can pick up to serve the last backup. This is of course expensive for any provider for every customer. But asking DO to actually report the status of the node, its host machine and the region is the right thing to do.
Customers don't need to know the full technical detail but even a nice friendly message (email, sms or even on the status page) will ease the conflict: "Your host now appears offline because the host machine is offline. Don't worry! Your data is safe with our backup! If you have any concern, please contact XXXXX@digitalocean.com or at xxx-xxx-xxx."
Conclusion:
* report the status of the droplet on personal dashboard
* for non-isolated incident, report it on both droplet personal dashboard and public dashboard.
For example, with Dreamhost you can get a VPS for ~$15 a month. If your e-mail isn't worth that much to you, then why are you bothering to self-host your e-mail anyways?
Because Google Apps no longer offers free accounts?
It should take some engineering work, but not whole lot.
How is that demand too much? Should we discard that demand because DO is a low-budget VPS? If you truly value your customer, you would take that suggestion seriously. I don't have millions to employ someone to manage an AWS farm for me. Instead of me asking DO why my nodes are down every time that happens, I want DO to tell me once that happen. It's a simple customer demand.
That being said, their penalties are ridiculous. I got a 10$ "SLA Credit" for 2 days of downtime. So I agree their SLA is useless.
Fixed that for you.
This should not surprise you. This is a common failure method of VMs. So, let's say a host was down. Depending on their storage methods, this means that the all the images on the disk are inaccessible. This means that you can't interact with it, which is why you couldn't take snapshots.
> Had to hammer their support with a dozen ticket before someone didn't just give me a canned reply.
Abusing support is never the right answer, I'm not surprised you only got canned replies.
> I like DO but stuff like this just can't happen without anyone checking on it stat.
It's almost like every problem cannot be solved instantly.
First it may be a $10 or a $20 box. But more importantly, every VPS provider of the few I've tried my self (e.g. Linode, XenCon), sends an email and opens a support ticket every time a $10 VPS goes off.
So please, do not try to change the norm. A server is a server no matter the cost and should be reliable. I like digital ocean but they won't get better by petting them.
Unless you're paying for an SLA which defines uptime and notification time, you're at the beck and call of Best Effort support, which may mean no notification and no reimbursement for downtime. Notification within minutes of your VPS going down isn't necessarily the norm.
Digitalocean stands to make ~$10,000 monthly of of a "$1,000 server" because it is somewhat managed, and thus has the responsability of informing the 100 or 1,000 odd customers affected because of their mistake imho.
Well no, I'm not paying commodity pricing (I'd hardly call a ~$6K server commodity) and OVH offers managed virtual rack services.
>> Digitalocean stands to make ~$10,000 monthly of of a "$1,000 server"
Err, what? You're vastly underestimating the cost of one of their nodes.
If my VPS is down and my hosting provider do not aknowledge it, it is only logical that I will have myself to create a ticket. As a customer I get both anxious and lose time. On the provider's end, now they have an open ticket that they have to read, possibly connect with another employee than the one that provides 1st level support and give me a proper answer. An automated ticket that would probably be semi-automatically closed too is best for both of us.
For notification purposes there are excellent services (uptimerobot, pingdom, monitor.us). SLA would be nice to have but most good lower cost VPS providers usually are over 99.9%, which is nice except the occasional outlier.
This is 100% a tool problem, and a problem we're actively working on for customers of ours that plan on having thousands of customized "views" for what they normally consider a "status page". Per-user functionality is one use case, but it can and will go deeper than that. If they could post an incident such that only you can see it, or such that only you are notified, they most certainly would.
I disagree that posting everything to be globally viewable is the right course of action, as this outage doesn't necessarily implicate fault on DO as a provider, but it also doesn't mean that you as an individual customer shouldn't have access to your specific view of a status page as it relates to exactly what infrastructure you live on.
You'd be surprised how prevalent this issue is, and how much inaction it creates on the provider end.
Most people don't go to a status page to check out how their provider as a whole is running. They just care about stuff that is directly relevant to them.
The whole point of IaaS is that the hardware/hypervisor/network is the provider's responsibility, and everything inside of your VM is yours. If the provider isn't doing due diligence to monitor their infra, then what infra are you selling as a part of your IaaS?
I think it's completely unreasonable to expect a status update on every single thing that might go wrong in Digital Ocean's infrastructure. If a single hypervisor/server fails, that could be ANYTHING. Bad drives, flaky memory, failed fans, etc., etc. This stuff happens ALL THE TIME and does not warrant a system wide update that something is wrong with the service, simply because there is nothing wrong with the service. All the other bits are functioning normally.
A loss of a box is expected and shit happens. Architect for it, or expect it to fail at some point. Everything dies eventually.
Also, $5 a month.
THIS is the power of virtualization, and if you're not doing this, you're doing it wrong.
So, I'll get a callback from my provider when my server is down...
Unfortunately, that callback is POSTed to a server that is already running on my provider, so I never see it :)
(Assuming that you mean that the callback is an external thing, i.e. provider -> customer. Also, I suppose I could host the callback receiving server on another provider. But I don't want to.)
Sometimes you actually do get what you pay for.
Otoh when they have whole api outages for creating or destroying vms (like also happened last week) i'd expect and did find a status update. What their threshold is for reporting isn't clear.. but a 1 machine issue isn't something any provider would report on.
The people that want HA like this generally custom build every server and don't use configuration management.
I'm hoping to finally build a tool this year incorporating some network monitoring (i.e. "unconfirmed reports") and a copy of all our internal rack/network database, so that we get to the holy grail: 1) customers get notified by email/SMS of stuff that definitely affects them, 2) it's easy for an engineer (or the whole team) to write notes on outages as they happen, and have them presented in a way that doesn't confuse customers.
This is the most important point. Status page should be a _log_ rather than a transient message as with so many providers.
edit: as other people have already namedropped, I'll point out that OVH do a pretty good job with the status page[1] and network maps[2].
One startup I know was using Heroku to host their website. One day the website had problems and was unavailable for extended period of time. Heroku team worked hard to resolve the problem, but it still took them several hours. Heroku did not update their status. They stated that if 99% of customers do not have problems, they consider any problems to be local and not reflecting status of the whole infrastructure.
The flexibility of the cloud is awesome but sometimes it does get cloudy up there. Redundancy is a must with digitalocean, aws, or anyone else.
My point was more that the frequency at which I and others are having DO issues seems to be higher than AWS recently (I can't independently verify it until I've used AWS myself) and truth be told it's a subjective/opinionated statement. I do remember a time when AWS was extremely flaky when they first started, so I'd like to think it's merely growing pains for DO, but at the same time ... I have my own growing pains to worry about.
I'm currently using round-robin DNS for load-distribution as the simplest thing that could possibly work, but it obviously doesn't actually balance load and it doesn't remove dead servers from the pool. What's the next step up that doesn't cost and arm and a leg to implement?
Don't expect your instance to run forever. Design your architecture with this in mind. There is a reason e.g. aws provides multiple availability zones.
DO droplets do have more disk space for your dollar, but AWS's answer to bulk data storage is S3, which is very cheap itself.
The main point is that S3 egress costs are substantial for bandwidth-intensive applications.
I take your point, though this being said, it again comes down to use case. A couple of hundred dollars a month for bandwidth is nothing for a business above a certain size, but it's tons for personal use. Depending on your use case, it may be also far more effective to use S3 and simply pay the data bill, rather than architect a system distributed across a ton of droplets.
I've been with Digital Ocean for over a year, and I'm fairly certain that I haven't had any downtime at all. The site is used by thousands of users a day, and I've never had any complaints. Pingdom is set to a resolution of 1 minute, and hasn't reported any outages either (that weren't caused by me).
Looking at the columns at a glance, there are almost no events "in progress" (most are Closed) and the vast majority of the open events are early warning for maintenance windows affecting very specific services.
As fellow geeks I am sure we have all been asked for advice on buying computers from our family and friends. I’m a mac person and use a $3000 laptop but I know most of them will not be willing to spend this much money. I normally quote them a $600(ish) computer that has at least an i-Series processor and 6-8GB of RAM. I tell them to avoid a couple vendors that I consider to be bottom of the barrel. It never fails though, that for all the reasoning and advice I give them, if they find a computer for $400 in the Sunday paper all that gets thrown out the window. They don’t want a computer that fits their needs they just want the cheapest computer. It also never fails that the moment it starts having problems they come to me for help and if I give them any slack for buying a cheap POS that I am suddenly considered an a-hole for not helping since that’s all they can afford.
The way I look at it is I spend my hard earned money on a reliable machine so I am not put out by the type of issues you get with cheap hardware or cheap services and you did not head my advice to avoid the same issues and thus by asking me to give up my weekend or even a couple hours to help you makes you the a-hole not me.
I have used, Rackspace, Linode, DO, AWS, and Google Apps among others and none of their status pages are every very helpful. It’s really a problem with Google Apps since my users know to check there and then claim the problem is not with Google even though it is. I frequently have issues with their IMAP servers failing where a user can connect via the web interface but not through a IMAP client. This is never shown on their status page. Of course I am going to check the status page.
The only hosting provider I have no complaints about is Rackspace but those servers are almost $1000/mo. On the other hand they do open up tickets for me faster than I can log into the management portal to do it myself. Even still I have had hour-long outages. If you don’t have HA you WILL have outages no mater who you use or how much you pay. Ironically the worst service was from The Planet even though those where still $800/mo dedicated servers. Their status page was a twitter account that they did not advertise on their homepage.
I have debated moving our average SMB size clients with no HA from linode to DO just because tools like Packer can interface with them easier. Still don’t know about that.
I have only been with DO for a short time but I have already encountered several of these silent outages where I can't query the status of my server through their tools and there is no status update. It is very different to Linode. They are a fantastic service for personal sites but I wouldn't host anything professional with them until they do more to gain my confidence.
Low prices and popularity bring growth.
Growth brings more customers than you can support as well as infrastructure problems that you can't handle. Because your company has not scaled up over a long time it has to hire people and shove things through the pipeline quickly which means mistakes (both in process and people) will inevitably be made. Rome wasn't built in a day as the saying goes but startups are. And they end up growing quicker than they should. Someone has to lose and it's the customer (not all of them but some of them). [1]
Not to mention the fact that if you are charging very little ($5 per month is pretty cheap obviously for the base service) it gives you less profit to handle things in a way that are perhaps more robust or doesn't give you the ability to paper over problems by building in redundancy.
The saying "price quality speed" pick any two applies here.
DO will get better of course but it will take time as they iron out and encounter the various issues that they face.
[1] I've observed this since 1982 when PC clones came out and competed with IBM. The clones shoved things into the channel and all the sudden hardware problems were shoved on the customers. Previously IBM charged enough that those things were handled by IBM not their customers. Because they had the profits and took the time to pay attention to details.
They should have a private/client-only per-hypervisor status page.
Exelion.net is running a special for HN users. 50% off for life for any server. Or more than one server. Or quite a few servers. Just use HN50 when checking out.
Exelion does top of the line E3-1230v3 Haswell (8 thread) + 16gb + 2x240GB SSD RAID 1 + Gigabit port + 33TB/mo bandwidth for $115/mo.
DO does 8 thread + 16gb + 160gb SSD + 6tb/transfer for $160/mo plus another $1350/mo for the full 33TB/mo bandwidth.
And if you wanted unmetered gigabit? Exelion does it for $200/mo extra flat rate. DO does it for an equivalent of $16,120/mo for that plan (based on 5 cents per gigabyte and 6TB already included with the plan).
It just doesn't make sense financially to keep using them.
servers fail, and you need redundancy. you might as well get angry at the sky for being blue.