Heroku is down
heroku.com
heroku.com
I know most sites hosted on heroku are pet projects with no real need that type of uptime, but this type of downtime makes it impossible to use them as enterprise customers.
Heroku had been a joy to work with until now. I ask of you all (as most of you are far more experienced in these matters than I), what is the traditional practice to mitigate this sort of risk? Paying for hosting with two separate companies?
EDIT: I understand the benefits of cloud hosting over hiring a sysadmin. At the same time, I'm interested to learn about what possible solutions there are. I'm at the point where I don't know what I don't know, and even the name of a topic or technology would be a huge help.
My apps triggered errors left and right and I was bashing my head against the keyboard because I noticed the errors earlier than their status update and I believed it was our fault.
Plus this last week the latency of the requests have been incredibly high, while the traffic to my apps have not increased. Again I assumed it was my fault.
I'm contemplating moving back to EC2 or Linode.
We moved all but one of our Rails apps off of Heroku precisely because of the frequent downtime -- or, rather, that was the last straw; there were other issues, notably the difficulty in debugging production issues, that had us already debating such. Heroku has gotten somewhat better, but it's still down far more than anything else that we use. (And we have services spread across Linode, Rackspace, AWS and the mentioned one app on Heroku.)
You'll have better uptime in most cases with a standard nginx / passenger setup on a $20 VPS than you will with Heroku.
I'd seriously consider Heroku if they were in multiple regions (fuck AZs, those are a lie). There are huge advantages to a PaaS in terms of speed and lack of hassle, and a well run PaaS is better than most developers at operations, patching, security, etc. (We're kind of unique in that we're better at operations than development, though.)
The competition for something like Heroku is EC2, VPS, or dedicated servers, in commercial colocation facilities. A good hosting facility is going to be a lot closer to 99.995% uptime for network and power to the box, but you can of course screw up past that point on your own.
99.9% is very doable. If your sites have less then you should consider hiring better staff.
Also the heroku figure (99.97%) is a lie - just skim their status-page.
In over four months we've probably already had more downtime than in the past 5 years. Despite paying quite a lot for "redundant", clustered offerings. There was one nasty bug in the hypervisor that killed the entirety of one of our providers' vps infrastructure - for over a day.
Our experience with "the cloud" was a little bit better, but we aren't entirely satisfied.
I don't know, maybe we should just order services from three different providers mirror our applications ourselves as a fail-over mechanism.
Yes, yes a thousand times yes. And dedicated servers don't cost much more than VPS's.
Perspective: I could throw an app on Rackspace and Amazon, with a replicated database, for under $200, in about a day.
For example, I really like AWS's security groups and ELBs. Those serve as my firewall and my load balancer and SSL terminator.
Replicating the application to another service means configuring and testing all that on my own.
If I use heroku and use their logging system, then replicating it to another provider means I need to be an rsyslog expert.
I don't really want to be an expert on rsyslog, postgresql configuration, floating IPs for HA LB, the best IO scheduler for file systems, etc. As someone who is in charge of all the sysadmin duties, and is solely responsible for writing all the business and db logic for several e-commerce sites, I want to spend my time on writing code. Not fucking around with figuring out the syntax for iptables.
The solution is either to have a PaaS provider who ruthlessly eliminates single points of failure (there isn't one, currently), or use some standardized software system which can be operated by multiple independent operators with nothing shared. Unfortunately, the only vendor-independent infrastructure is the physical server, various forms of VPS, etc. -- it's all at the IaaS level. As far as I know there's no PaaS type thing with a common interface which arbitrary providers can operate, with some kind of marketplace for users to pick operators independently from the technology.
This isn't possible by definition, right? The PaaS provider itself becomes the single point of failure.
Rings a bit like a "Who created God?" argument. "What single entity can I use to defend against failures by a single entity?"
You can mitigate specific risks, and you try to prioritize those based on cost, frequency, and severity. If there were a great redundant provider with good authentication on accounts, a strong balance sheet and business, and sane policies on managing accounts, you would be fairly safe using just that provider. After all, you could always get a court order to cease providing services, yourselves, like if you do something some troll has patented. It kind of depends on your application, too -- if I were doing a wikileaks, a bitcoin exchange or torrent site or some other legally at risk business, I'd want country-level separation across multiple providers, at least as a cold backup. Casual game for facebook or mobile, not really much of a concern.
logging
security hardening
firewalls
backups (both point in time and complete backups, both for database and for any other artifacts like image uploads)
testing backup recovery (both the point in time and complete backups)
replication
HA
LB
SSL termination
postgresql configuration
monitoring (security, system stats, application availability, individual process running)
performance analysis (both at application and system level)
method for deploying updates (to the application and to the system)
and the list goes on... you could easily make a career out of focusing on postgresql configuration, for example. these aren't simple things to do.
Different providers have different limitations on what you can do here. For example, AWS doesn't support multicast, which limits your options for doing High Availability.
Almost all of the things you outlined are things that have in the past blown up in real world deployments, often at cloud hosting providers where you didn't know about it it because a well-paid, well-trained ops team hid the drama for you.
So yeah, definitely factor that in to the cost of hosting at a cloud provider. You're absolutely right.
It's not an insurmountable barrier, for sure. Knowing how to operate your own code is a great feeling.
B) A multihomed setup is not a black box of mystery, and is nowhere near as expensive as people are saying from the hip.
Then "you," having acquired years of experience in designing and implementing distributed services, are making at least a 6-figure investment in opportunity cost.
That is to say, I know my way around Linux, have set up single/standalone servers for all sorts of services, but have never known remotely where to begin on the distributed side of things, much less understand what is required to makes something "designed to be multihomed".
Fun? Yes. Miss it? Fuck no. Recommend people relive the 2012 version of the experience? Huh? Get back to work writing code. Re-read the Wikipedia page on "comparative advantage" first if you need to motivate yourself.
An experienced ops team willing to be on call 24x7 is easily six figures by itself.
With that in mind, everything I talk about in my reply is doable by one person, or a solo developer. A cursory understanding of administration goes far, and if you're deploying to Heroku, you should know enough about administration to not be completely in the dark when things fall apart.
Deploying a mass-consumable Web service implies that you have accepted the fact that you are now a developer and administrator.
When I get such a screen I immediately close the website I'm visiting. And no, contrary to what they tell me, I do not have "a virus" nor is my computer part of a botnet.
customer=> Cloudflare SSL
then
cloudflare=> Heroku SSL through Heroku Piggyback
https://openshift.redhat.com/community/blogs/new-openshift-r...
Obviously it's our fault for not having a redundant system / for trusting heroku. Though I still can't help but be a little pissed as I email 100 people about how sorry I am for their service disruption. Heroku didn't email me.
On the other hand its surprising to me that when heroku goes down, it goes down hard, namely that all hosted sites including their own are unavailable.
This is probably harder then it seems, especially when the outage is related to their routing infrastructure.
"We have confirmed widespread errors on the platform. Our engineers are continuing to investigate."
Edit: Also IRC #heroku
Edit 2: heh... I just noticed the "subscribe to notifiations" link on the incident page: https://status.heroku.com/incidents/372
- it would be possible to have the non-DB part of your heroku app be spread across multiple availability zones.
- worker pricing would be much lower or based on actual CPU cycles used.
- add-on providers would be vetted more thoroughly before getting a spot in the add-on store.
- they'd reopen the #heroku irc channel for informal support
At least then, it won't be "Heroku is down", it will be "Heroku West is down" or some such thing.
Last time I can recall this happening to us it was down for several hours.
See: http://en.wikipedia.org/wiki/High_availability#Percentage_ca... for a handy chart and follow along with any of your favorite *aaS providers!
just a glitch in the matrix. it occurs when they switch over local control.
Edit: Came back for a second. Went down again. Came back. Went down.
Potential Platform Issues 5m+ We have confirmed widespread errors on the platform. Our engineers are continuing to investigate.
this is the message for Production and Development
Potential Platform Issues Jun 7, 2012 15:55 UTC - Update
We have confirmed widespread errors on the platform. Our engineers are continuing to investigate. Posted Jun 7, 2012 15:58 UTC Issue
Our automated systems have detected potential platform errors. We are investigating. Posted Jun 7, 2012 15:55 UTC