"DigitalOcean Killed Our Company"
twitter.com
twitter.com
This situation occurred due to false positives triggered by our internal fraud and abuse systems. While these situations are rare, they do happen, and we take every effort to get customers back online as quickly as possible. In this particular scenario, we were slow to respond and had missteps in handling the false positive. This led the user to be locked out for an extended period of time. We apologize for our mistake and will share more details in our public postmortem.
That said, what's highly troubling as a DO customer (and someone who is planning to deploy startup infrastructure of my own with DO) is:
1) The discrepancy between this customer's experience and clear assurances made on this very forum by high-level DO employees that:
a. warnings are ALWAYS issued before suspensions.
b. even in the event of a suspension, services remain accessible (though dashboard access and/or the ability to spin up NEW services may be impacted), ie. the affected customer could still retrieve data or SSH in to droplets.
2) The relatively trivial nature of the customer's offending usage (temporarily spinning up 10 droplets). What happens if, for example, a startup gets a press mention somewhere that leads to a massive traffic spike, necessitating a sudden and significant spin-up of new droplets (especially if this is done programmatically versus by hand in the dashboard)?
3) The apparent lack of consideration of the customer's history, or investigation into their usage. It seems the threshold for suspending services of longstanding customers who are verifiably engaging in commerce (taking a moment to look at their website and general online presence for indicators of legitimacy), should be SUBSTANTIALLY higher than for, say, an account who signed up a week ago. Context matters.
Following is a comment[1] by Moisey Uretsky in another thread[2]:
> Depending on which items are flagged the account is put into a locked state, which means that access is limited. However, the droplets for that account and other services are not affected at all. The account is also notified about the action and a dialogue is opened, to determine what the situation is. There is no sudden loss of service. There is no loss of service without communication. If after multiple rounds of communication it is determined that the account is fraudulent, even then there is no loss of service that isn't communicated well in advance of the situation.
1. https://news.ycombinator.com/item?id=18296344
2. https://news.ycombinator.com/item?id=18294940
This is why I'm so confused by the case under discussion, because the customer appeared to have been completely locked out without warning.
If DO reserves the right to cut off services and access to your own data permanently and without warning (outside of a court order or confirmed illegal activity), that needs to be unequivocally stated, and the triggering factors should be made known. Otherwise, DO is not fit for production systems.
Additionally, it would be nice to see the creation of a transparent, high-level appeal process for customers affected by suspensions. Truly malicious customers wouldn't use it (what would they hope to successfully argue to an actual human reviewer?), but it would greatly benefit legitimate customers to have an outlet other than social media by which to "get something done" in the event of an inappropriate suspension followed by a breakdown in the standard review process.
This didn't seem like a case of being "too slow" - the customer in question went through your review process (which was slow, yes), and the only response he got was "We have decided not to reactivate your account, have a nice day".
That just seems like a lack of interest in supporting your customers that are falsely flagged.
Hope you can share what you learnt from this incident and hopefully you'll take a hard look at your processes.
I'd hate to be caught in the same issue, especially that we are already customers, and I'm not sure I'll have as much clout as Nicolas here to get your attention.
It's occurring to me now that while I've successfully ignored twitter for years, I should probably rectify that just so I have somewhere to type my hopes and prayers when this eventually happens to me, and hope for a miracle. It sure seems like the only place they're listened to.
Maybe keeping a twitter (and other social media) account with at least a certain number of follower should be considered a part of a company's security strategy? You'd also need to post something interesting periodically, to keep your follower, so that you have their attention when you need it.
https://news.ycombinator.com/item?id=6983097
Running anything business or privacy critical on DO is madness.
It's fair to note that scrubbing is now the default behavior when a droplet is destroyed, so they did listen to the feedback.
You do not need to scrub or write anything to not provide user A’s data to user B in a multi-tenant environment. Sparse allocation can easily return nulls to a reader even while the underlying block storage still contains the old data.
They were just incompetent.
On top of all of that, when I pointed out that what they were doing was absolute amateur hour clownshoes, they oscillated between telling me it was a design decision working as intended (and that it was fine for me to publicize it), and that I was an irresponsible discloser by sharing a vulnerability.
Then they made a blog post lying about how they hadn’t leaked data when they had.
Nope.
Companies are made of people. Let the people have a life. Their night is shitty enough as is after this, I guarantee you.
If you have ever been involved in post facto analysis of a process breakdown like this you know how hard it is to get the full picture immediately. Rushing something out does no one any favors.
"This will not happen again, ever".
People's livelihoods are at stake in DO's hosting. Canned responses and brutal account lockouts should have NEVER been on the table to begin with.
That’s akin to saying “we’ll never ship a bug”, or “we have an SLO of 100%”. That’s impossible for anyone to claim. Same goes for the response handling. There is clearly a lot of room for improvement there, but if you’re insisting on not getting canned response, that means a human needs to be involved at some point. Humans will at times be slow to respond. Humans will at times make mistakes. This is just an unavoidable reality.
I get that mob mentality is strong when shit hits the fan publicly, but have a bit of empathy and think about what reasonable solutions you may come up with if you were to be in their situation, rather than asking for a “magic bullet”.
I could see a good response here being an overhaul of their incident response policy, especially in terms of L1 support. Probably by beefing up the L2 staffing, and escalating issues more often and more quickly. L2 support is generally product engineers rather than dedicated support staff/contractors, so it’s more expensive to do for sure, but having engineers closer to the “front line” in responding to issues closes the loop better for integrating fixes into the product, and identifying erroneous behavior more quickly.
However, can you say with a straight face that the very generic message left here by DO's CTO instills confidence in you about how will they handle such situations in the future?
Techies hate lawyer/corporate weasel talk. Least that person could do was do their best to speak plainly without promising the sky and the moon.
I’m an engineering manager in an infrastructure team (not at all affiliated with Digital Ocean, tho full disclosure, I do have one droplet for my personal website). I know how postmortems generally work, and it’s messy enough to track down root cause even when it’s not some complex algorithm like fraud detection going off the rails.
I’d rather get slow information than misinformation, but I understand the frustration in not being able to see the inner working of how an incident is being handled.
And I agree with your premise. However, my practice has shown that postmortems are watered-down evasive PR talk, many times.
If you look at this through the eyes of a potential startup CTO, wouldn't you be worried about the lack of transparency?
And finally, why is such an abrupt account lockdown even on the table, at all? You can't claim you are doing your best when it's very obvious that you are just leaving your customers at the mercy of very crude algorithms -- and those, let's be clear on that, could have been created without ever locking down an account without a human approval at the final step.
What I'm saying is that even at this early stage when we know almost nothing, it's evident that this CTO here is not being sincere. It seems DO just wants to maximally automate away support, to their customers' detriment.
Whatever the postmortem ends up being it still won't change the above.
I too value less known providers. The human factor in support is priceless.
Do you believe that a PR response made in damage control mode that actually changes nothing is something that's satisfactory?
I mean, apparently this screwup was so damaging that it killed a company. What part of the PR statement addresses that precendent?
But the DO CTO did basically admit fault in a public forum.
0) https://www.digitalocean.com/legal/terms-of-service-agreemen...
Additionally, the steps taken in our response to the false positive did not follow our typical process. As part of our investigation, we are looking into our process and how we responded so we can improve upon this moving forward.
As a business owner with much of our infrastructure depending on DigitalOcean, the incident is concerning. It affects the reputation of DO as well as its customers.
The demographics on Twitter and especially here on HN represents a sizable crowd with decision-making influence on DO's bottom line. I hope to see some effort being made to prevent situations like this in the future, and to regain the trust.
As a (so far) satisfied customer, it's great to hear that:
> A combination of factors, not just the usage patterns, led to the initial flag.
> We recognize and embrace our customers ability to spin up highly variable workloads, which would normally not lead to any issues.
> we are looking into our process and how we responded so we can improve upon this
The real “mess up” here was the bit where you blocked the account with no reason given and no further communication - other than the one-liner your intern wrote for the email.
I’m expecting you to sit down with your legal team and rewrite your TOS to be more customer-focused and less robotic.
Looking forward to the write up!
I know people say some legal arguments why they close you down and won't say anything, but this is the worst scenario ever. I'd be better off excused at something I didn't do than just ooops we can't tell you anything, your account has been shut down.
The response email even read like a giant polite FUCK YOU (we locked your account, no further action required by you)
You bet I will have further action!
And it is after the shaming that you get an "I am sorry for this situation". Which sounds more like saying "I'm sorry we got caught".
My frustration is not with DO specifically, as they do exactly what every other company does.
But, what of the other thousands of people that got screwed and did not put it on twitter?
It is the equivalent of when you are in a restaurant and get screwed: It is the loudest person that complains more the one that gets the reward, while all the others silently swallow the injustice.
Getting your message into the right hands is what matters, not the platform it's on.
Considering you have been marketing yourself as the platform for developer oriented cloud, you should be aware that surge provisioning can and will always be happening.
But it doesn't make sense to shut it down before discussion!
The very fact that this can happen from an automated script with no oversight should give every one of your customers pause as to whether they continue with your service.
I block their netblocks for a lot of things, too.
That doesn't sound good to me.
Ideally you should be cloud-agnostic, but that's quite hard to achieve.
It requires so many failures in understanding the service being provided across the company for this decision making process to have ever actualized that there is no reasonable expectation of safety or trust from DO at this point.
You clearly don't make every effort, and did not -- so why waste the extra verbiage and switch from active to passive voice?
Based on your cliche response I have zero confidence that DO will do anything substantial to address the root causes of the issue.
I'd gladly do whatever it takes to KYC, send you my business license, tax returns, EIN, invoice billing, etc so you know there is someone behind my account.
We spend thousands of hours eliminating single points of failure. If an automated system can undermine that work, DO is not an option for us to host anymore.
Yesterday we completed our postmortem analysis of the incident involving Nicolas (@w3Nicolas) and his company Raisup (@raisupcom). With their permission we are sharing the full report on our blog here:
https://blog.digitalocean.com/an-update-on-last-weeks-custom...
DO's tier 1 support is almost useless. I set up a new account with them recently for a droplet that needed to be well separated from the rest of my infrastructure, and ran into a confusing error message that was preventing it from getting set up. I sent out a support request, and a while later, over an hour I think, I got an equally unhelpful support response back.
Things got cleared up by taking it to Twitter, where their social media support folks have got a great ground game going, but I really don't want to have to rely on Twitter support for critical stuff.
DO seems to have gone with the "hire cheap overseas support that almost but doesn't quite understand English" strategy, whereas the tier 1 guys at Linode have on occasion demonstrated more Linux systems administration expertise than I've got.
They told me that on a single day a support engineer was supposed to help/advice customers on pretty much whatever the customer was having issue with and also handle something between 80-120 tickets per day.
It's nice to see that DO is willing to help on pretty much anything they (read: their team) has knowledge about, but with 80-120 tickets per day I cannot expect to give meaningful help.
Needed EDIT: it seems to me that this comments is receiving more attention than it probably deserves, and I feel it's worth clarifying some things:
1. I decided not to move forward with the interview as I was not interested in that support position, so I have not verified that's the volume of tickets.
2. From their description of tickets, such tickets can be anything from "I cannot get apache2 to run" to "how can I get this linucs thing to run Outlook?" (/s) to "my whole company that runs on DO is stuck because you locked my account".
Disclaimer: Based on providing support myself and coaching a support team at both a web hosting company and ISP I used to own years ago.
you are entirely correct.
Gotta hit that ZBB!
Now its just "meh, we'll fix it in next months release".
Anyway imho you should have taken the support position and schemed your way into development internally. This was my plan at eBay before they fired me, though they shut down the branch here a few months later and moved to the Philippines anyway so I wouldn't have lasted long regardless.
Let me stress here, this is not nearly as easy of a problem to solve as it appears to be on the surface. We're struggling as a company right now because after our recent merge, a lot of our good talent has left and we're having to rebuild a lot of our teams. Even so, I'm still happy with our general approach. Management understands that employees will often have wildly different problem solving approaches and matching metrics, and that's perfectly OK as long as folks aren't genuinely slacking off and we as a team are still getting our customers taken care of. I think that's important to keep in mind no matter how big or small your support floor gets.
I couldn't imagine getting that level of support from DO, let alone Amazon.
Even when we had small handful of physical servers with them, they seemed inept. They actually lost our servers one time and couldn’t get someone out to reset power on our firewall.
That said, our FAWS team are a good bunch, and what AWS lacks in support they more than make up for in well engineered, stable infrastructure. Since Rackspace's whole focus is support, I think the pairing works well on paper and it should scale effectively, but we'll have to see how it plays out in practice.
https://www.rackspace.com/managed-aws
This is a big push, internally and externally. I don't know too much about the details (I don't work directly with that team) but it's been one of our bigger talking points for a while now.
It's crazy that companies spend $$ on marketing and sales, then cheap out on a interaction with someone who is already interested in / using their product.
"We have a mark, lets suck it dry until we can throw it away and find new and better marks."
Sustainability tends not be a concern until after a group's leadership jumps ship and parasitizes other hosts.
Almost invariably the high ticket rates are also driven by bad product elsewhere. Money is being spent on customer "services" sending out useless cut-and-paste answers to tickets to make up for money not spent on UX and software engineering that would prevent many of those tickets being raised. Over time that's the same money, but now the customer is also unhappy. Go ask your marketing people what customer acquisition costs. That's what you're throwing away every time you make a customer angry enough to quit. Ouch.
> This was my plan at eBay before they fired me
I guess you answered yourself.
The reality is that his boss is true...
I somewhat blame people in tech, actually. More than one company is creating products that "cut customer service costs via machine learning", which is code for "pick keywords from incoming tickets and autoreply with a template"
But the execution matters a lot, and DO's is currently not great. IIRC, it takes clicking through a few screens of "are you sure your question isn't in our generic documentation? How about this page? No? This one then? Still no? You're really sure you need to talk to someone about this error? sigh Okay, fine then."
These systems should not be implemented as a barrier to reaching human support, but they often are.
I truly feel sorry for their first and second tier customer support people. I imagine the staff churn rate is incredible.
People who work for these sorts of low-end hosting companies inevitably quit and try to work for an ISP that has more clueful customers. When you have people paying $250/month to colocate a few 1RU servers, the level of clue of the customer and amount of hassle you will get from the customer is a great deal less than a $15/month VPS customer.
I look at companies selling $5 to $15/month VPS services and try to figure out how many customers they need to be set up for monthly recurring services, in order to pay for reasonably reliable and redundant infrastructure, and the math just doesn't pencil out without:
a) massive oversubscription
b) near full automation of support, neglect of actual customer issues, callous indifference caused by overworked first tier support
Conversely, as a customer, you should be suspicious when some company is offering a ridiculous amount of RAM, disk space and "unlimited 1000 Mbps!" at a cheap price. You should expect that there will be basically no support, it might have two nines of uptime, you're responsible for doing all your own offsite backups, etc.
If you use such a service for anything that you would consider "production", you need to design your entire configuration of the OS and daemons/software on the VM with one thing in mind: The VM might disappear completely at any time, arbitrarily, and not come back, and any efforts to resolve a situation through customer support will be futile.
I can tell you that as a person whose job title includes "network engineer", we have a number of customers who have critical server/VM functions similar to these people who had the DigitalOcean disaster. If something goes wrong, an actual live human being with at least a moderate degree of linux+neteng clue is going to take a look at their ticket, personally address it, and go through our escalation path if needed.
There seems to be a sweet spot for company size here. Too small companies can't support you even though they really want to. Large companies are busy chasing millions in big contracts, and don't really care about your $800 per month at all.
If you are just some semi anonymous faceless person ordering services off a credit card payment form on a website, all bets are off...
That's going to be true no matter which cloud provider you choose.
Their ToS almost certainly include terms which allow them to kick you off and refund any monies for any reason whatsoever.
Good luck if you bought into their entire ecosystem and can't move elsewhere on a whim.
My company provides services to fortune 100 companies, and we host literally petabytes of data on their behalf in Amazon S3, but we don't have offsite backups. We (and they) rely on Amazon's durability promise.
We do offer the option of replicating their data to another cloud provider, but few customers use that service -- few companies want to pay over twice the cost of storage for a backup they should never need to use when the provider promises 99.999999999% durability.
Reasoning: Your contract with Amazon promises durability and I'm sure there's a service level agreement with penalty/liability clauses. By implementing a redundant backup, you're replicating something that you don't legally need to have, double-or-more due diligence on the offsite backup security/credentials, and in case of a failure of Amazon create a grey area with clients "Do you have the data, or do you not?"
In short, there could be a very good business reason not to do offsite backups.
"We're sorry, the tape that we didn't needed to keep has been lost/zero-dayed/secondary service provider has gone bankrupt/Billy's house that we left it at got robbed." These must be disclosed to a customer immediately.
Minimising attack/liability surface is not only a technical problem, but a business one too.
I'd expect that, as more transistors have been packed blah blah blah, that such a $10/mo account would have gotten better, not worse, since then.
This is what Linode do - they keep you on the same payment level, but raise what you get for it.
I think that the size of the company or how fast they grow is a good proxy for having poor customer support. What we should be doing is finding the slow growing businesses or the mid-tier (not too small, not too big) businesses to take our business to.
Although these days I tend to go with UpCloud, it is very similar to Linode and DO, except you can do custom instances like 20x vCPU with 1GB Memory, spin it up for $0.23 an hour. Compared to standard plan on Linode and DO, 20 vCPU with 96GB Memory would be $0.72/Hr.
https://hn.algolia.com/?query=linode%20security&sort=byPopul...
They've had several high-profile breaches over the years.
A Bitcoin theft via the Linode Manager interface in 2012: https://news.ycombinator.com/item?id=3654110
and a second Linode Manager compromise in 2016: https://news.ycombinator.com/item?id=10845170
In both cases, if you dig into the context a bit, the story turned out that Linode wasn't fully disclosing the breaches to their customers until they were forced to do so when the news about them reached a certain volume. They also may have been -- almost certainly were -- dishonest about the extent of the damage and how it may have impacted their other customers.
At the time, their Manager interface was a ColdFusion application, which tends to be a big pile of bad juju. They started writing a new one from scratch after, I think, the second compromise.
The really bad thing here is that they got soundly spanked for being less than truthful the first time, and then four years later -- when they'd had ample time to learn from that mistake -- they did it again.
So there's a nonzero chance at any given time that Linode's infrastructure has been compromised and they know it and have decided not to tell you about it.
That's what prompted me to start exploring DigitalOcean more. Unfortunately, I've found that there's a far greater chance that I'll experience actual trouble exacerbated by poor support than that I'll be impacted by an upstream breach, so about half my stuff still lives on in Linode.
This is entirely the roots of my distrust of them right now. Mistakes happen. Companies I trust demonstrate that they've learned from their mistakes. My tolerance for mistakes is pretty low when it comes to security related things, though. If something has gone wrong, let me know so I can take remedial steps. Their handling of both of those incidents suggested I can't trust them to tell me in sufficient time to protect myself.
Just my experience though.
I suspect as they've increased in popularity they've become a bigger target for DDOS attacks.
I've also noticed that in the past year there's been a lot of data centre outages... like every couple months. Hasn't been a deal breaker for us, since our traffic is generally fairly low outside of season, but the ones during seasons really hurt.
Also I'd like to add that they do give you the heads up when there are issues, which is a big plus in my book compared to some other hosts.
I really do think it's just growing pains, and I don't mean to disparage them. Just being honest that I wouldn't recommend them for high availability services. Since I consider them a low budget host that's probably unfair though. We've just outgrown them is all.
This is all anecdotal of course.
That's simply not true. There's support engineers hired around the world, and depending on when your ticket is posted, someone awake at that time will answer. DO is super remote friendly and as a result, has employees (and support folks) everywhere on the planet. Not "cheap oversea support" at all. There's a lot of support folks in the US, Canada, Europe, and in India where they have a datacenter.
disclosure: ex-employee
Going to guess from this tweet https://twitter.com/AntoineGrondin/status/113096281882239385... that you're currently staffed at DO. That's fine; I know two people who do great work on DO's security team. But it would be helpful if you could disclose this when you comment about your employer publicly so that readers don't have to dig up your keybase and then your twitter account to understand it.
Then of course, there is no guarantee these people speak and understand english perfectly.
The reality is that top talent, even if remote, is competitive whether they are in NYC/SFO/SEA or not. And DO has some pretty talented people on staff.
And then, having people in all timezones is definitely an advantage for 24/7 support. I'd say it's not negligible, and not an after thought.
Now about english fluency, it's only that important to english native locations. And really, most of tech does not necessarily have english as a first language - I certainly don't. So I'd say that encountering support engineers with imperfect english shouldn't be a problem to anyone, and definitely not a sign of cheap labor. In fact, I'd say bitching about someone's english proficiency in tech is kind of counterproductive and I find it discriminatory.
Anyways. DO doesn't hire international employees to get cheap labor, that's a preposterous proposition. And with datacenters around the world and a large presence and customer base, it makes sense to have staff on board from many of these areas. And that staff might answer your tickets at night when they're on shift. Shouldn't they?
That said, I disagree wholeheartedly that it's okay for support staff to not be completely fluent in the language they're providing support in, regardless of the language.
There is functionally no difference between trying to interact with talented support staff who aren't fluent in your language, and trying to interact with illiterate support staff. The end results are identical.
There are people who are very talented and very fluent in more than one language. Those people tend to be more expensive. So, many companies forego hiring those workers and instead hire others who are cheaper and "about as good". My multiple experiences with DO support have suggested that that's what they're doing.
As other commenters are suggesting, it may just be instead that DO is expecting its support staff to meet metrics that are causing them to spend only a minute or two per ticket and send out scripted replies.
People who would make gratuitous grammatical mistakes but have read more classics than the average American college graduate. I can easily count many just thinking of it.
> There is functionally no difference between trying to interact with talented support staff who aren't fluent in your language, and trying to interact with illiterate support staff.
That statement reeks of ignorance. It seems you have almost no experience with other languages than your own, or you would know that communicating while being non-fluent or with a non-fluent works just fine most of the time. Sometimes misunderstandings happens and it can be a bit slower but that is all.
fluency is a high barrier to clear. it took me 5 years of speaking/reading/writing english daily to come even close to "fluency".
before that, i had a really good advanced english, but i wasn't fluent. and it didn't mean i was "illiterate".
So your position is that support that's a bit slower, with some misunderstandings, is exactly as good as fast support without misunderstandings, even in downtime-sensitive applications.
Well, okay then.
isn't a native English speaker, or speaks in different dialect doesn't mean that they are any less capable
It's not just about being less capable... Being less understandable can trump capability.Linode even has an irc channel(you can use browser to access it), I rarely need support, but when I really need it, it is always fast, to the point, available.
I was eager to try the new DO managed Postgres service, but I guess I won't after this blunder.
DigitalOcean should go the route of AWS and kill off free support completely and offer paid support plans. Something like $49 a month or 10% of the accounts monthly spend.
If you are a serious company with paying Fortune 500 customers, you need to act serious and pay up for premium support and stop expecting free.
You already get bumped to a higher tier of support once you hit 500/month spend for no additional fees but that just gives you promises for lower response times
It's usually a bad sign when a company can't meet current needs/expectations and then decides to try and productize their failures.
Main, potential caveat is they're Xen-based hosting. That may or may not matter depending on what one wants to run. They support the major Linux's.
Can't comment about support, but DO's tutorials on linux server random task 101 are fantastic.
I was surprised how much of my ubuntu server setup googling ended up on DO pages.
The difference in support DO vs Linode is probably due to DO being cheaper.
What? Most DO and Linode plans have exact same specs and cost exactly the same, and IIRC it was DO matching Linode. Although DO didn't seem to enforce their egress budget while Linode does.
(I've been a customer of both for many years, and only dropped DO a few months ago.)
EDIT: Also, according to some benchmarks (IIRC), for the exact same specs Linode usually has an edge in performance.
I've had interesting issues automating deployment to the point that my current build script provisions 9 VMs, benchmarks them and shuts down the worst performing 6. Some of their co-location is CPU stressed.
When there's an issue with Linode's platform, they discover it before I do and open a ticket to let me know they're working on it. When there's an issue with Vultr, the burden of proof seems to be on me to convince them that it's their problem not mine.
That said, any company, especially one working with Fortune 500's, should have DB backups in at least two places. If they'd had the data, they could have spun up their service on a different hosting provider relatively easily.
I find it concerning that they have such a low threshold for triggering a lock. 10 droplets is hardly cloud scale.
-When the majority of abuse support dealt with was people angrily calling and asking about the fraudulent charges on their cards for dozens of Lie-nodes you consider putting caps in place to reduce support burden and reduce chargebacks.
At the time at Linode, if it was a known customer, we could easily and quickly raise that limit and life is good.
I've always wondered how Amazon dealt with fraud/abuse at their scale.
I don't think DO was wrong here to have a lock, but the post lock procedure seemed to be the problem.
Even worse: they don't explicitly state what is considered "unreasonable". So, if your business is serious, you have to assume the worst-case scenario: DO can't be used for anything
Conclusion: Digital Ocean is just for testing, playing around, not suitable for production.
I think that's always been the standard position most people take. DO, Linode, etc are for personal side projects, hosting community websites, forums etc. They are not for running a real business on. Some people do, sure but if hosting cost is really that big a portion of your total budget you probably don't have a real business model yet anyway.
That text suggests larger organizational problems within the company.
Yes they should.
How many 2-man shops do you think follow all the proper backup and security procedures?
Every week there's another article on HN about a tiny business being squished in the gears of a giant, automated platform. In some cases like app stores this is unavoidable, but there are plenty of hosting providers to choose from. People need to learn that this is something that can happen to you in today's world, and take reasonable steps to prepare for it.
You can't dot every I, cross every t and also build a compelling product as a 2 person shop.
How many horror stories need to reach the front page of HN before people stop believing this? Getting locked out of your cloud provider is a very common failure mode, with catastrophic effects if you haven't planned for it. To my mind, it should be the first scenario in your disaster recovery plan.
Dumping everything to B2 is trivially easy, trivially cheap and gives you substantial protection against total data loss. It also gives you a workable plan for scenarios that might cause a major outage like "we got cut off because of a billing snafu" or "the CTO lost his YubiKey".
Sounds like the opposite of the survivor bias. I don't believe it's any sort of common (though it does happen), even less that "it should be the first scenario in your disaster recovery plan"
I could not disagree more. There's a right way and a wrong way to do this, it's trivial to do it right, and the risks of doing it wrong are enormous.
Then it's unrealistic to trust them with your business.
We don't know the structure of their DB and whether failover is important or not, so we don't know if the DB can be reliably pulled as a flat file backup and still have consistent data.
We also don't know how big the dataset is or how often it changes. Sometimes "backup over your home cable connection" just isn't practical.
Cron jobs can (and do) silently fail in all kinds of annoying and idiotic ways.
And as most of us are all too painfully aware, sometimes you make less-than-ideal decisions when faced with a long pipeline of customer bug reports and feature requests, vs. addressing the potential situation that could sink you but has like a 1 in 10,000 chance of happening any given day.
But yes, granted that as a quick stop-gap solution it's better than nothing.
I'm going to take a stab at small and infrequently.
Every 2-3 months we had to execute a python script that takes 1s on all our data (500k rows), to make it faster we execute it in parallel on multiple droplets ~10 that we set up only for this pipeline and shut down once it’s done.
Even if they lost some data, even if the backup silently failed and hadn't been running for two months, it's the difference between a large inconvenience and literally your whole business disappearing.
They had backups, but being arbitrarily cut-off from their hosting provider wasn't part of their threat model.
Isn't a big part of cloud marketing the idea that they're so good at redundancy, etc. that you don't need to attempt that stuff on your own? The idea that you have to spread your infrastructure across multiple cloud hosting providers, while smart, removes a lot of the appeal of using them at all. In any case, it's also probably too much infrastructure cost for a 2-man company.
keeping your production and your backups in the same cloud provider is the equivalent of keeping your backup tapes right next to the computer they're backing up. you're exposing them both to strongly correlated risks. you've just changed those risks from "fire, water, theft" to "provider, incompetence, security breach"
...
> > "fire, water, theft"
i'm sure you could add a few more things to the list.
> I don’t think it’s terribly common for even medium sized companies to have a multi tier1 cloud backup strategy.
not terribly common to understand risk.
> So what is the purpose of the massive level of redundancy that you are already paying for when you store a file on S3?
You're paying to try and ensure you don't need to restore from backups. Our data lives in an RDS cluster (where we pay for read replicas to try and make sure we don't need to restore from backups) and in S3 (where we pay for durable storage to try and make sure we don't need to restore from backups), but none of that is a backup!
If you're not on the AWS cloud S3 is a decent place to store your backups of course, but storing your backups on S3 when you're already on AWS is, at best, negligent, while treating the durability of S3 as a form of backups is simply absurd.
> I don’t think it’s terribly common for even medium sized companies to have a multi tier1 cloud backup strategy.
The company I work for is on the AWS cloud, so we store our backups on B2 instead. It's no more work than storing them on S3, and it means we still have our data in the event that we, for whatever reason, lose access to the data we have in S3. Who the hell doesn't have offsite backups?
This is not remotely the same thing. A RAID offers no protection against logical corruption from an erroneous script or even something as simple as running a truncate on the wrong table. Having a backup of your database in a different storage medium on the same cloud provider protects from vastly more failure modes.
> Who the hell doesn't have offsite backups?
No one. But S3 is already storing your data in three different data centers even if you have a single bucket in one region, and you also have SQL log replication to another region. Multi-region is as easy as enabling replication but that is only available within a single cloud provider (I can't replicate RDS to Google Cloud SQL, only to another RDS region). I would guess that a lot of people use that rather than using a different cloud provider.
That sounds like...the same argument?
A RAID array stores your data on multiple physical drives in the machine, but offers no protection against logical corruption (where you store the same bad data on every drive), destruction of the machine, or loss of access to the machine.
S3 stores your data in multiple physical data centres in the region, but offers no protection against logical corruption, downtime of the entire region, or loss of access to the cloud.
You can't count replicas as providing durability against any threat that will apply equally to all the replicas.
Let's be fair: The threat model here is "lose access to our data".
This can happen in a number of ways, lost (or worse, leaked) password to the cloud provider, provider goes bankrupt, developer gets hacked, and a thousand other things.
Even if you trust your provider to have good uptime, there's really no excuse for not having any backups. Especially not if you're doing business with Fortune 500's.
Literally just push a postgres dump to S3 (or any other storage provider) once a night as a "just in case something stupid happens with my primary cloud provider". It'd take a couple hours tops to set up and cost next to nothing.
Also, by "two places" I meant the live DB and one backup that's somewhere completely different. My wording may have been confusing.
Containers solve the easy problem, which is how to make sure the dev environment matches the production environment. That is it.
Replicating TBs worth of data and making sure the replica is relatively up to date is the hard part. So is fail over and fail back. Basically everything but running the code/service/app, which is the part containers solve.
> Sure, a backup would have been a significant improvement, but still – a backup only protects against data loss and not against downtime.
Assuming you have data backup / recovery good to go, the downtime issue needs to be solved by getting your actual web application / logic up and running again. With something like docker-compose, you can do this on practically any provider with a couple of commands. Frontend, backend, load-balancer -- you name it, all in one command.
> Containers solve the easy problem, which is how to make sure the dev environment matches the production environment. That is it.
Speaking of "patently false"...
> Account should be re-activated - need to look deeper into the way this was handled. It shouldn't have taken this long to get the account back up, and also for it not be flagged a second time.
So... he doesn't address what is the scariest part to me, the message that just says "Nope, we've decided never to give your account back, it's gone, the end."
I think it's entirely reasonable for companies to have that option. "You are doing something malicious and against the rules, you have been permanently removed". In this case, that option was misused, but I don't think the existence of that possiblity is inheritly surprising.
And, regardless of what DO should or should not do, they can do whatever they want with their own hard drives. You should structure your business accordingly.
And do what in the mean time? The legal system acts slowly. In the age of social media outrage, would you allow the headline "Digital Ocean knew they were serving criminals, and they didn't stop them" if you were CEO?
It's easy to be outraged when these systems and procedures are used against the innocent. That does not mean we should stop using rational thought. If someone is using DO to cause harm, then DO should (be allowed to) stop the harmful actions.
You lock down the image, and let law enforcement do their thing. If law enforcement clear them, you then give the customer access to their data, perhaps for a short time before you cut them off as they seem to be a risky customer to have.
You don't unilaterally make the decision, you offload your responsibility onto the legal process.
The fact that there are hundreds of comments on HN condemning them for this action proves my point.
Seems to work just fine for AWS, Google and Cloudflare. In fact, counter to your argument, Cloudflare got in massive shit when they did decide to play God.
The observant will note the particular corner you're backing into here -- that a business might be justified in denying access to code/data being used in literally criminal behavior -- is notably distinct from the general and likely much more common case.
> they can do whatever they want with their own hard drives.
Sure. But to the extent they take that approach, Digital Ocean or any other service is publicly declaring that however affordable they may be for prototyping, they're unsuitable for reliable applications.
Businesses that can be relied on generally instead offer terms of service and processes that don't really allow them to act arbitrarily.
I agree. Look at the absolutism of the comment I am replying to. My whole point is that there might be some nuance to the situation.
> ...Digital Ocean or any other service is publicly declaring that however affordable they may be for prototyping, they're unsuitable for reliable applications.
Again, I agree. Considering how cheap AWS, backblaze, and Google drive is, it is completely ridiculous to depend on any one single hosting service to hold all your data forever and never err.
You seem to be accusing the aggrieved party of being a bad actor, when that is not the case.
If there is no police report, then they are trying to act as police themselves, which I think is unacceptable. It is not their data.
Your argument that they can do whatever they want with their hard drives is indeed something I will take care to remember — I definitely would not want to host anything with DO.
At the very least, they should also provide ALL, as in every last byte, of data, schemas, code, setup etc. to the defenestrated customers. As in: "sorry, we cannot restart your account, but you can download a full backup of your system as of it's last running configuration here: -location xyz-, and all previous backups are available here: -location pdq-".
Anything less is simply malicious destruction of a customer's property.
If you violate a lease and get evicted, they don't keep your furniture & equipment unless you abandon it.
They should have, at the very least, one DR site on a different provider in a different region that is replicated in real-time and ready to go live after an outage is confirmed by the IT Operations team (or automatically depending on what services are being run).
Support was useless and even with evidence did not believe who I was.
I then somehow convinced them to give me temp access, which in my opinion is even worse. They didn't believe me about who I was and then gave me temporary access to an account. DO can't be trusted when their support team could so easily be socially engineered.
What???
Unfortunately, my ticket no longer appears under closed tickets. I was still able to dig up my original ticket message and all the responses their support made to me through my email though. Here they are:
Between the replies I asked about what I could do to verify my account. As you can see, they didn't even give me a single chance to do so. They told me to hang on twice then just permanently closed it up. I'm not sure how I even got flagged. All I did was turn on a droplet and delete it. I checked the audit logs, and there was nothing suspicious there either. It was just me logging in and out.
I thought about making a deal about it on Twitter, but I didn't bother because I don't have any followers and it wasn't a huge loss to me either. Maybe it's the only way?
"We've locked you out, no explanation" "Sorry for any inconvenience"
Seriously? That last line is like a slap in the face.
No one should talk to a customer like that in this situation, if only because (a) if this is real abuse, you don't need to be "fake nice" and (b) if it's a false-positive, you've just come across as extremely smug when you're in the wrong.
Wow. F*ck you too.
"We've tried nothing and we are all out of ideas! LOCKED" is basically their response.
I'd complain on Twitter just to see what happens.
I've recently implemented backups with Restic, the static binary and plethora of supported storage backends was extremely appealing. The easiest seems to be to just point it at a S3 bucket, but given most people have infrastructure on AWS (off-premise means off-premise) having other options supported out of the box is pretty handy.
Sure, having other options is good: https://min.io/
On the other hand, completely shutting down all their services without a quick conversation...
I know of a company that explicitly had a "call us to discuss first" clause in their contract with a smaller cloud provider. Everyone was on holiday and not answering the phones while their hacked account was being used to spin up dozens of boxes launching a DoS attack against a crypto scam site. Guess who had to eat the bill on that one?
Granted that would require people to communicate and use some form of reason.
Even DMCA for all its warts fires up a warning and has a response mechanism (granted other issues there).
I'm sure many people have started their companies firmly convinced that they'll give plenty of warnings and never automatically shut anything down.
The problem is, you rapidly discover that doesn't scale, not even on a human level. You send your notice. 48 hours later, you've gotten no response. If you act now, it isn't materially different from your point of view as if you simply acted right away.
Also, in a cloud environment, even Digital Ocean, as many people have learned the hard way with leaked credentials, you can rack up charges faster than the relevant humans can even conceivably be notified. As the hosting company, you can't just let abusive or accidental usage go. You can refund their money, but that's still resources of yours that went to something that failed to produce revenue rather than something that did; you can't absorb that indefinitely.
I'm pretty sure you'll inevitably discover that you have no choice but to put automation in.
It does probably help that I said I would be careful not to do that again and had already put in a CloudWatch Alarm to automatically power-off the instance after a set period of idleness before filing the ticket.
I will gladly pay the extra money for AWS than to even think about DO or even GCP for a money making project.
But more on topic: with Aurora/MySQL you can have an on-site hosted read replica from an AWS hosted database. That would be a cheap, easy real time backup solution if I were really worried about AWS screwing me over.
The proposition seems to go something like this: it's a new thing, mistakes are statistically expected, you make an honest one and plead "oops!" and we refund you, no doubt pointing you to resources on best practices and account throttling. As long as the customer takes the lesson to heart, everyone wins.
The host should throttle resources in the interim if its at risk of running up massive bills or adversely impacting other clients. None of this is breaking new ground, there isn't a good reason for large hosts to act like shit.
If I’m a $10/month customer, kill my account early, and it’ll save me more often than not. If I’m a big spender, maybe wait a bit longer.
No warning was given, and no way to retrieve any data. Fortunately nothing essential was lost.
A fly by night will often only suspend the VM or database that is in question, not other services on the account (having been in that position before).
Hardly fly-by-night.
The one time I used their phone support, the guy I got was fairly helpful. They seem like a hands off company overall.
Good practice, actually, but it's something people usually get frustrated with and they're literally famous for doing it.
Edit: I stand corrected ¯\_(ツ)_/¯
And they also have a lot of experience with it because they host a lot of game servers which are very prone to DDoS attacks (see e.g. https://securityaffairs.co/wordpress/51640/cyber-crime/tbps-...), so they are definitely “structured to offer it”.
On OVH one of my servers was hacked and running typical scripts that are run once that happens (port checking, common admin credentials, brute force attempts, etc)
They cut off all internet access to and from the server and sent me an alert stating what was happening and that I needed to VNC into the server, resolve the issue, and let them know how/why it happened and how I resolved the issue. Once that was done they just removed all the blocks on the server and we all went on our merry way.
Edit: To clarify the VNC console is on their site, not a remote connection.
I didn't know what they were talking about so I replied saying that, they helped me shut down the local port connections and I never heard another complaint from them. There was no downtime or banning at any point.
I saw a great deal on LEB for a KVM VPS from Alpharacks and signed up for a 2 year plan (my first mistake).
When SSHing in to the VPS it didn't have the advertised specs, and when I raised the issue with their support, they eventually fixed it..
Then I realized the second problem, they gave me the same IP address as someone else. You could still use the web VNC console, and as soon as you made an outbound network connection, inbound connections would work... for a few seconds... then SSH would drop. Reconnecting by SSH says "host key changed" i.e. you hit someone else's server sharing the same IP address. Using the web VNC console, works again for a few seconds, drops again.
It took about 7 days of arguing with their support to explain these two problems to them... by which time, the 3-day refund window had expired...
I admit some schadenfreude watching their recent disaster (all servers down since the last 2 weeks - https://www.lowendtalk.com/discussion/157613/popcorn-time-du... ).
Lesson learned about low-end boxes. I have had about 50/50 good and bad experiences, using half a dozen providers like this, not really making financial sense overall.
I recommend everyone stay far far away from AlphaRacks. If anything remains of them after this week.
Developer has a Python script that takes 1 second per record to execute and he has 500,00 records to process, so he spins up 10 distinct VMs each running the same Python script to parallelize the task.
The provider shuts him down and cites a section of the EULA that says "You shall not take any action that imposes an unreasonable... load on our infrastructure." Basically saying "Hey whatever you did, don't do that."
Developer gets his account restored and then proceeds to do exactly what he did to get it locked out the first time around.
Also Developer has all his eggs in one basket.
Shitty customer/product service aside, someone explain to me why DigitalOcean is at fault here?
> to make it faster we execute it in parallel on multiple droplets ~10 that we set up only for this pipeline and shut down once it’s done.
Tildes preceding numbers means "approximately". Why doesn't he know exactly how many VMs he spun up?
Was he actually spinning up 1 VM per record and only allowing 10 VMs to be running concurrently?
I'm not a Pythong dev but why can't you execute 10 instances of Python on a single VM?
If you need to dedicate an entire VM to processing 1 row, what the hell is it doing?
Random speculation of one possibility: each of those 10 instances were suddenly doing something unexpected and spammy with the network. Maybe sending 500k+ emails (one per row of data claimed by the developer) over SMTP in a very short period, or jumping to massive spikes in torrent traffic, or crawling sites to scrape data (maybe each row of 500k is just a top-level domain name, and they crawl every URL on those domains, possibly turning 500k rows into hundreds of million of http requests).
The postmortem will be interesting. If DO is truly at fault here, that email after the second lockout saying the account is locked after review, no further details required... bad.
Additionally, if 10 spun up VMs is considered an “unreasonable” load on DigitalOcean infrastructure I shutter at the concept of building anything on the service. Does DigitalOcean even define “unreasonable” in their terms or is it kept vague?
Developer explains as requested.
Developer receives "k, we've restored your account", which sounds an awful lot like "what you're doing is fine".
Developer gets shut down again despite having explained the behavior as requested (i.e. the explanation is on file) and despite the explanation being considered sufficient by DO.
The "with the details you provided, we've removed the hold" e-mail is hard to interpret differently than "sorry for the misunderstanding, this is fine", especially as the initial e-mail asked them to explain to "ensure your account is not subjected to additional scrutiny or placed on billing hold". If they meant "ok but don't do that again" they should have stated so.
Think of that automated monitoring like a circuit breaker in an electrical panel. If you plug an air compressor into a 15 amp circuit and it pops the breaker, once you flip the breaker back on would you run the compressor again? No? Then why would you think it ok perform the same activity that tripped up the automated system?
The OP seemed to be aware that spooling up 10 VMs to do whatever he did was what got him booted, so... why didn't he reach out to DO to find out exactly what alarms he tripped and then take action to either get his account whitelisted or modify his process to ensure it didn't trip the automated alarms again?
He just got his account unbanned (flipped the circuit breaker) and fired his job back up (turned the compressor back on).
(And you cant just start 10 hugely expensive VMs either, the larger sizes are initially locked too)
Are we getting to a point yet where that’s considered to be suicidal? I hope so.
* Autobans in facebook
* Cheated instacart drivers
* $10000 stolen from thousands of bank accounts (and returned hopefully) on Etsy
* Tesla cars literally killing people (it now feels like it's once a month right?)
Now the software that runs software is running amok.The interesting thing about software is that it runs very quickly and it acts as a giant lever that affects the entire world.
You can think of it like a giant airport suddenly being installed in your back yard and just start having planes take off, changing your $300k investment into a $120k valued house overnight. That's how quickly software is changing the real world.
I know there is at least one HN browser writing a book on it. But I would love to see more books on how the internet, and software, is messing our world up.
It's quite scary how low our standards have gotten.
I really don't understand this sheeple thinking, for example most people simply don't understand that an automated fraud detection system is not a technologically important thing, it's an economically important cost management system effectively.
Just like companies externalize the cost of helpdesk personnel by operating an automated call center (and by proxy making the customer bear the cost), the goal is the same with fraud systems. But we cannot just simply throw our arms up in the air or shrug our shoulders when the companies leverage our lives this way.
Take Facebook for example, they acted like they had no responsibility or any power to review and take down or prevent toxic and/or hateful comments by employing human reviewers until they were forced to do so in some countries. And guess what, they had no trouble doing so, their profit might have reduced somewhat, but not that much.
So all in all, anyone who thinks he has integrity as a developer should take a look into himself when he justifies systems like fraud systems (or any other unnecessary cost reducing actions) as necessary. They are economically beneficial, yes; necessary, no.
That's what destroys company reputation.
I may be wrong but my understanding is that the gold standard - Amazon Web Services - will only ever suspend your account until an issue is sorted out.
Whoever runs Digital Ocean needs to stand up and say very loud and clear to this community that they will never, ever delete accounts - if he doesn't then he can live with the business destroying reputation of Digital Ocean being "one of those account killing companies".
What company would ever host on "an account killer" - the risk is way too high.
Because fraud and abuse exist.
Sometimes customers really are doing malicious things which need to be stopped immediately, and the only thing that will make that happen is disabling their account. Trying to make accommodations for those users is a fast path to getting sued, getting blacklisted by mail providers, and/or losing your upstream connection entirely.
At the very least this company learned a hard lesson about Disaster Recovery best practices. Hopefully all the up and coming companies reading this story learns as well. Also please remember that a backup that isn't tested IS NOT A BACKUP! I've been in so many situations where backups were corrupted, so part of the disaster recovery is to test the backups and make sure you can really recover.
There are a million different scenarios where their data could be lost and it not be Digital Ocean's fault. It's the company's responsibility to have protected their customers from this.
I worked at a company with no real Disaster Recovery plan. I was told that "we can get the servers up and running within 18 hrs if we had an outage", which not only was absurdly slow but probably an underestimate. Only by the grace of God did we not suffer a real outage but if we did, it was totally the VP's fault for not addressing my concerns.
Often the supplier has to fill out the checklist themselves.
> do you have backups?
Manager: [X]
> are they offsite?
Dev: cloud provider docs say so
Manager: [X]
Customer: Great, you won the bid.
It's a shame to see things play out this way, but sometimes a lesson is taught in a brutal fashion.
I imagine the author will be more careful in the future regarding off-site backups, additional technical partners, contingency plans, etc.
I'm not talking 5 9's redundancy. I'm talking grab a backup once a week or something, anything, to help mitigate a scenario like this. According to the thread, they lost ~1 year of data. That should be unfeasible to a company serving customers, let alone Fortune 500 customers.
Disaster recovery planning is key for a technical company to succeed. It is clear they never considered a scenario where there DO account would be closed/compromised/down.
I don't think the chance of DigitalOcean automatically freezing your account to a point where only a co-founder can do something about it has been well publicised.
A contingency plan should ideally have been in place for a scenario where, regardless of root cause, you have lost access to your DO account.
Given their size, that is extremely unlikely to happen without warning.
Imagine you are a customer of this company. Would you be rallying to their defense, "backups aren't needed because the scenarios are unlikely", or would you be angry that the company had zero contingency planning and lost all of your data (or the data you rely upon)?
If you can honestly say, as a (hypothetical) customer of the company in the thread, that you wouldn't care if a company you relied upon has no disaster recovery planning, more power to you. I, however, like to make sure that the companies I'm relying on have some sort of contingency that protects me as a customer.
An expensive lesson to learn to take and check backups regularly.
Stuff like this is relatively simple when you are trying to learn it, but keeping it operation is hard at a small level. Is the less to really use PaaS until it's viable for you to be running a small K8s cluster, or equivalent fleet? Seems really expensive compared to VPSes, but having better guides might help.
It won't be entirely up to date when the worst case happens, you'll be unavailable and you'll probably have lost a day of data or so, but you won't have lost everything.
Rclone to AWS or Google or whatever is easy to set up, add a daily dump of your database to the folder you back up. Unless you handle a lot of data, costs are probably not a big factor.
Clearly they should be charging more.
Even the database replication is mostly setup and forget these days. Very reliable.
And they already did backups. It's just they did them to the wrong (same service) host.
Yes, it sucks that DO did this. But this is hardly the first time someone got screwed over by some poor AI automated security. Backing up your data to backblaze or AWS would be the cheapest insurance policy you could buy.
I find this enforced mediocrity pretty appalling. With barely functional "anomaly detection" Deep Learning models with dubious decision making (I did some so I am familiar with the "landscape") it's gonna be a lot of fun for anything slightly deviating from whatever vague norm that can't be explained nor tested against.
Wait, a whole distributed computing sub-system to make a 1s process faster?
> I got their final message right after arriving in Portugal.
Did he initiate the script from a IP address outside France?
Though that might explain why DO's fraud detection was triggered, it doesn't excuse their actions. Send an email first, jeez.
I think OP means 1s per row, so 138 hours.
This is a shame, and imho it sucks a lot. One of my biggest points of paranoia are with backups and scripting a return of a site on another provider should the worst happen.
Getting ready to launch something barely more than a hobby and was planning on DO because the hosted Postgres and a small K8s cluster is significantly less for what's there than the alternatives. Frankly, I don't want to go from ~$100/month to over $200 for another provider for something that likely won't lead anywhere.
Goosfrabah... goosfrabah...
I think what you have to explain is why there wasn't a contingency plan, with your own servers, colocation, another cloud offering, etc...
Why would there be an expectation that a 2-man shop have "another cloud offering" as a contingency plan when some of the biggest and best tech companies do not?
People use services like AWS or DO because they are the contingency plan - they have the size and scale that smaller companies cannot afford or implement.
I'd argue that it should be _easier_ for a 2-man company to adapt to cloud service outages, as they likely don't have to keep up with nearly as many backups or moving parts.
Which would mean you disagree strongly with coldtea?
Their entire business was completely reliant on DO droplets. It doesn't take much foresight to think, "hey, I should probably make a backup in case this VPS goes down."
Nothing in this comment thread, or the OP twitter thread mentions anything about the rest of this imaginary contingency plan of theirs.
PostPost said they didn't, that even huge companies don't have contingency plans.
I agree with PostPost, and I'm trying to figure out which one you agree with.
If you define being able to adapt as a contingency plan, well, I have confidence that this company is fully able to adapt! Their architecture is small and pretty easy to move. The only problem is a lack of external backup, which will be remedied very soon, and once that happens they could easily shift to another service even if DO re-disabled their account.
So that would mean you agree with PostPost. But you don't seem to agree at all.
I'm struggling to reconcile "The ability to adapt is the definition of a contingency plan." and "this imaginary contingency plan of theirs". If you demand a preexisting written plan then that means you're not accepting "the ability to adapt" as a valid answer at all.
Yes, ideally they would have already tested their back-up solution, the back-ups would be offsite and, if something like this happened, they could stand-up on another provider. But that just ignores the reality of them being a super tiny business. Almost no one at that size is going to do that.
But I still support that DO here is a 100% liable toward their client. Now the liability between the said client and his own clients is an other matter.
Ideally partners should be trusted (and trustful), in practice, they aren't
Though trusting DO/AWS/GCP, etc is much more reliable, than, let's say, betting your whole business on somebody's proprietary API (like an FB game, an Linkedin API something, etc)
> Digital Ocean "Trust and Safety"
Does this phrase give anyone else the heebie-jeebies? They deliberately locked his account without looking at previous metrics over the previous months to conclude this was obviously not a concern, and why?
Things that have crossed my mind:
Is their automatic system was poorly tuned? Was this deliberately initiated? If so, why?
What happened w/ Digital Ocean is inexcusable, and has potentially dire consequences for two individuals livlihood. In the immediate aftermath of such an event, focusing on the devs percieved lack of disaster preparedness seems petty.
- No offsite backups? Agreed. Even for a two person team it is sloppy.
- "Relies on one tech partner?" Strongly disagree. Even large enterprises often have a hard dependency on AWS, Azure, Rackspace, or similar. To suggest that a two person team should have deployment plans for multiple independent cloud vendors is just fantastical thinking with no basis in reality.
Plus if they did over-engineer it by making it cloud agnostic, setting up accounts to sit dormant, cold instances elsewhere, etc, people would just criticize them for that inefficiency/wasting time.
How you got from "single tech partner" to "have no disaster recovery plan" I don't know.
Very, very few tech companies could simply move everything to a new cloud provider in a few hours. I would even hazard a guess that almost none can.
I have all my infrastructure as code and can break it all down and spin it back up in kubernetes clusters in minutes. But due to the quirks of each cloud provider, there are tons of little fixes that would inevitably need to be made.
Not to mention that many companies have way more data than could even be copied over in a few hours.
I think we needlessly shame one-person operations for focusing on actual customer needs instead of ops busywork and yak shaving.
No one that I've seen so far is saying they should have a system that is "muilti-site, fault-tolerant, self-healing, webscale that Google would be proud to have".
They are saying run a simple backup and keep it literally anywhere else.
I think we needlessly hyperbolize "do a backup once a month and keep it somewhere else" into some sort of NSA operation.
Its backups. Its 2019. Its dead-easy and very affordable.
The cofounder picked it up 3h ago. DO responded and apologized from the official account 2 hours ago, claiming it was fixed, and is actively responding to people tweeting at them, doing damage control, since about 1 hour, promising a public postmortem.
While it's sad that a social media escalation was needed (and it confirms that getting attention on social media is the only effective way to resolve hard issues like this), the response after that was quite fast. Let's see how well and how quickly they deliver on the postmortem.
Two years later I wanted to restore the VPS but turns out my snapshot become "outdated" and they stopped supporting the format for restoration... Support was completely useless, wouldn't even let me download the snapshot, said at most they could mount it into a new VPS and I could recover the data myself.
Very unprofessional.
I dug around to find a Wayback version that actually had text. I understand the words on their own, but when put together I get nothing: https://web.archive.org/web/20181030015237/https://raisup.co...
>powerfull
https://twitter.com/moiseyuretsky/status/1134547532149854208
yikes, just when i thought their kubernetes and managed db's were looking attractive...
While we could certainly survive the loss of these assets, the recovery would be long and costly.
So I would certainly say that this story gives me a great deal of pause and will take up some mental space this weekend as I think about future dealings with DO.
And yet, the explanation is very simple:
Because you neglected basic principles and elected to put all of your backup eggs in one basket.
Professional stuff is one thing, but that's not to mention anything personal - anything I care about I won't put exclusively on someone else's computer. I want to have absolute control over as much of my stack as I can. Seems really scary that some company has control of your entire infrastructure and can ban-hammer you without notice, permanently, at any time of day or night (or while you're on vacation).
Caps or quotas are not sufficient to deal with this problem.
It's not this anecdote in itself, but that it corroborates my experience during trial that I ignored and dismissed as support incompetence (which should have been a warning sign in itself). After setting up the account, adding a payment card, I wasn't able to enter our VAT ID as part of billing details, with some nondescript error. So I asked support.
Two days(!) later, they responded by asking for incorporation documents, which was frankly bizarre (and a first in ~10 years of running business): they're not exactly a bank with KYC requirements. When I responded, basically, WFT?, and told them to check the billing data in VIES, they eventually fixed it.
But what I got from it was a distinct impression that their default assumption is that the customer is trying to defraud them, even when it makes no sense. To this day, I have no idea what kind of fraud could they possibly be anticipating there (they allowed the card).
This story is on the same general subject, and so are others surfaced here and on twitter in reaction: the customer is presumed scumbag.
What I'm saying is: If you ARE going to put your business at the mercy of one company from top to bottom, would it not be wise to try to get some kind of account rep? Or have some sort of communication with the company as to the nature of your operation?
And if that isn't an option, should you do business with them?
I don't ever back up to the same service that I use for production. Or is that just me being paranoid?
Clearly not.
That's another reason that I'm very skeptical of Amazon and Google offerings of their proprietary API. Rent virtual servers, no problems, you can rent those servers from a thousands of hosters. But if you're using their proprietary APIs, you will have to rewrite your software to migrate and you will have very tight time frame.
I use 4 for my current company and have redundancy spread over them so that if, say, AWS goes down or I get an account locked or whatever, nothing is lost, things continue to be operational, slight degradation happens and that's it.
This really isn't that hard to set up. Under a day or so and then just do stuff in ssh config and the shell rc to act as helpers so you remember how to do things.
It's super cheap, pretty easy, and robust.
It's awful what happened to this guy but it's kinda like the person who backs up to the same harddrive as the originals. Awful to lose stuff, it shouldn't happen, but also don't do that.
I have heard although can’t confirm that using on account billing rather than using a credit card makes them less likely to just disable your account for billing issues. Things like that would be helpful to know. Should I be letting them know more contact info about us, asking for an account manager, etc?
Does it make a difference in how they react to you if you are spending low thousands a month vs tens of thousands a month vs hundreds of thousands etc? Or is everything always automated to death?
These sorts of things tend to slip through the cracks. If you are running is a business capacity, make sure you treat everything like any other business would.
Even as a two-man company, you can't afford the cost and potential liability of not operating as a registered business. It also shows that this is a serious business and not a hobby.
Is there any evidence for this? For all we know his DO account was set up as a business.
Not saying this shouldn't have happened, and hopefully they have learned from this experience.
Wait their entire business was effectively shutdown and all they did was send an email? Granted the DO handling of a possible abuse situation is shocking, but to allow your business to go down for 12 hours and not be trying to call any and every actual human being at DO seems to be negligent on their part. While us technical types love interfacing via digital means, some situations benefit from actually talking to a live human.
In the case of these business functions, the standard and correct approach is to outsource them to a 3rd-party service provider. You do this with accountancy, legal representation, facilities, office management, recruitment, etc. If you try to bring all these things in-house from the get-go you'll never get around to building a product, and it's financially and logistically sensible to do it with IT as well. This calculation may change over time as your business grows, but if you can't comprehend that the correct strategy for a fledgling business may not be the same as an established one, then you're simply not suited to run a business in the first place.
Of course, in every case, you're taking a calculated risk by relying on a service provider: They may go bust, they may be incompetent or malicious, they might ramp up their prices. Your job as the manager of a company is to accept and manage these risks as best you can. Risks cannot be eliminated, only managed, and attempting to do is a fool's errand. If things do go wrong, you'll always have people lining up to tell you how you could have avoided this problem, usually it's by making a decision that can be justified with the benefit of hindsight. You should ignore these people. The only question is: did you make the correct decision at the time, based on the facts to hand?
Also, had troubles with trial too - activate $5-10 coupon was quite impossible without PayPal payment.
Hopefully everything gets righted and these folks can start making solid, multi-site backups.
* The company quickly creates and starts 10 VMs and triggers an automated lock-down, with a message mentioning a sudden spike of activity.
* The company gets the account unlocked.
* The company again quickly creates and starts 10 VMs, and triggers the auto-lockdown again.
Note to self: when something damaging happens as a result of a seemingly normal action, avoid doing that same seemingly normal action immediately again, lest the damaging consequences hit again.
So the lesson is don’t leave your backups with the same cloud provider that host your database. You should also pull local copies as well
On the other hand - I'm becoming increasingly aware of the inevitable Twitter-social-media-pile-on for any company ever.
In these days, most apps generally can be migrated to a new host in seconds as long as you have the data source alive.
If they had access to thier data they probably should have been able to spin up a similar ec2 instance in minutes and say goodbye to DO forever.
Unfortunately it doesn't help that cloud egress bandwidth is criminally expensive.
One thing I will add is that, especially for a small shop or project, assume from the get go that by renting infrastructure from DO (or any provider) that user-hostile actions can and will be taken when it comes to any issue regarding TOS violations that you are unaware of.
This assumption helps to build redundancy in your mindset. Have a production website or app in DO for example? Droplet backup, periodic snapshots, B2 server backup, S3 tar backups, containerize apps if possible, have equally provisioned (smaller, idle VMs) infra on another DO account or another provider if possible, and so on. I know this is overkill but paranoid sysadmins/devops are always rewarded.
Just to add some context for DO specifically, they're a great provider in my anecdotal experience and they are constantly rolling out services aimed at medium to large scale workloads, such as managed databases and k8s.
That being said, it's entirely possible to transfer snapshots [0] to another DigitalOcean user account or teams account. So at the very least, create an entirely new DO account just for holding snapshots, outside of the native droplet backups and third-party backups you're doing on an application level.
[0] https://www.digitalocean.com/docs/images/snapshots/how-to/ch...
Say I have a friend at all the continents (including Antartica for hypthetical fun-ness), they are all willing to allocate some square meters to plunk down some servers for whatever is needed.
How would I go about it and build this small(ish) infrastructure myself?
DO in my experience is not the greatest company to work with; but at the same time, your "incompetence" killed your company, not DO.
https://revenni.com/cloud-contingency-when-the-ban-hammer-dr...
Even better: run your infrastructure across multiple hosting providers with something like consul. DO might cause slower service but not a death sentence.
This is why you need, at minimum, a nightly mirror. Ideally you stack on top of that a load-balancing device or service to redirect traffic in the event of an outage from your primary farm.
Never go on holiday again until you have backups and failover.
#idothisforaliving
https://revenni.com/cloud-contingency-when-the-ban-hammer-dr...
Our business is Dynalist, an online outliner app. Many of our users store all their notes on Dynalist, so uptime is really important.
Starting 7 PM last Tuesday, we saw a slowdown in request handling. We filed a ticket with DO 2 hours after that (we also posted our initial tweet to keep our users informed: https://twitter.com/DynalistHQ/status/1131087411797270529).
A few hours later, we started to experience full downtime. Still no reply from DO. We filed another ticket with the prefix "[URGENT]". Still no reply.
We waited for 24 hours for their reply. We took turns taking naps because we're only a 2-person team.
After 24 hours, we tweeted @ DO (https://twitter.com/DynalistHQ/status/1131397013306847232). 2 hours later we finally got a support person working on our ticket. We didn't want to take it to the social media, but there doesn't seem to be any other way at that point. DO doesn't have phone support, and us "bumping" our support ticket didn't work either.
After 2 hours going back and forth on the support ticket and providing logs, DO's support person identified the issue and offered to move us to a less crowded server. They asked us what's a good time to do a manual migration if a live migration fails, and we replied immediately saying whenever is fine (we're experiencing downtime anyway).
We thought it's over, but we were so wrong.
They didn't reply in another 4 hours. That was 4 hours of more downtime. Sometimes, CPU steal is down a bit and our server could catch up some requests, although it would still take 10 seconds for our users to open Dynalist. But most of the time, our web app was totally inaccessible. Watching the charts on our dashboard go up and down felt like some of the hardest hours of my life... mainly because there's nothing we could do.
4 hours in, I realized we had to post another angry tweet to get a solution. There's nothing else to do other than trying to stay awake anyway. So I posted another tweet: https://twitter.com/DynalistHQ/status/1131497962184564737
This tweet didn't seem to work. Nothing happened in the next 3.5 hours and things started to feel surreal. I didn't know how much longer this downtime is going to last, and I didn't know what we were going to do about it.
At that time, it was 9:30 AM EDT and people were starting their day. We were getting more and more emails and tweets asking what is going on and where are their notes. A few customers were angry, but most were understanding and supportive.
At 9:55 AM EDT, DO finally did the live migration a few minutes before the time limit we gave them, which was 10 AM. That was the end of the incident; CPU steal was down to < 1% and Dynalist was finally up again.
However, we couldn't trust DO any more. This weekend we're migrating to a dedicated server provider which has phone and live chat support. DO is pretty good for spinning up a $5 box quickly to test something, but we learned the hard way we shouldn't rely on it.
Our postmortem post: https://talk.dynalist.io/t/2019-05-22-dynalist-outage-post-m...
PS: Warheads have THREE copies of navigation systems
https://qbix.com/platform is one of many projects working to tackle this. Tim Berners-Lee’s SOLID project and others are, also.
The rise of the public cloud is pretty fortunate for companies like DO, they can make the same assumptions that legacy VPS companies did (most vm's sit idle so oversell them massively, most customers will create a single vm and nothing else) while branding themselves as having the same strong infrastructure as AWS.
IANAL
As far as sympathy goes, you’re not wrong. But you’re also justifying every pain in the ass procurement process you’ve ever dealt with. Your attitude is why so many companies won’t go near a two man shop.
They've basically been flagged as abusing the system multiple times, and they're surprised they had to kick up a storm to get themselves reactivated again?
Not to mention, that process they need to run every couple of months, that takes 1s, but they still need to parallelize over a bunch of vms, that's weird and sounds like something that needs to be rearchitected at the very least.
Just a guess.
Not at all trying to blame the original posters and victims:
While they seem great for hobbyist and small business sites, there’s no way I’d trust Fortune 500 client business to something like DigitalOcean. I just don’t see the benefits over a more established operation like AWS, Azure or GCP. Saving $50 here and there isn’t worth it.
You get what you pay for. We're even upgrading from this support plan to an Enterprise account.
Startups can usually get enough in AWS credits that they probably could have their entire first year of service _for free_.
Yes, this is a bad look for DO. But the way they're able to beat AWS on price includes things like "worse support." And if AWS goes out of business, you'll know in advance. DO isn't the same story. You should be planning for redundancy if DO is truly business critical for you.
(I'm not trying to blame author of that tweet, we all make mistakes and I think DO should give access to backups even when account is locked)
I think it is reasonable to expect your hosting provider won't randomly shut down your account...twice.
BTW, I don't even see a downvote button around posts, there is only upvote button! Is it because I'm new here?
Keep in mind though that you should only downvote for abusive comments (which also deserve a flag), gross misinformation, or other things like that. Disagreeing with someone is not a good reason to downvote.
And unfortunately this stuff happens to your "established" examples as well. Here [1] is a particular example of Google shutting down an entire GCP account with no explanation. Some comments report the same on AWS as well. Ironically, people in that HN thread are actually suggesting Vultr (kinda like DO, but even smaller) as a good alternative.
I'm lucky enough that my spend with DO is high enough to qualify for support, so if this ever happened to me at least I know I'd get a couple chances to make things right
Were this to happen on GCP I'm fairly certain they'd just black hole my account since I'm spare change to them.
After even a single incident like this no sane company would relay on DigitalOcean. This is the kind of crap you expect from a shared web host overselling resources, not from a company that wants to provide infrastructure to tech companies.
I don't think you realize how big is the price difference between AWS/GCP/Azure and old-school VPS providers like DO, Hetzner or OVH.
Regardless of this example of a false positive, locking whole accounts over that is unwise.
Also, even the co-founder doesn't seem to know exactly why the service was suspended, even though he clearly managed to arrange things.
So... Never trust DigitalOcean for anything important.
The response (and its timing) will determine whether we continue with DO as a host for the (admittedly tiny) bit of infrastructure we host there.
And there had better be a reasonable response on HN if they value their HN-reading customers; it is where we got the first recommendations for them years ago.
I'm referring to _this_ response, i.e., the one where they explain why they did what they did.