Heroku down?
status.heroku.com
status.heroku.com
"8:50 PM PDT We are investigating degraded performance for some volumes in a single AZ in the us-east-1 region."
AWS has been historically bad at reporting the severity of their outages promptly.
AWS is a fascinating science experiment. Pity about the websites, though.
God only knows what they reserve the red circle for...
In the past 30 days they have had 2 outages which have lasted more than 2 hours. That's a lot of down time.
Cloud hosting is fantastic but it's a trade off. There are so many layers of abstraction between you and the hardware that you are completely at the mercy of one, two or (more!) technical organizations, each with their own support systems and varying levels of opacity into their infrastructure.
The fact that Amazon went down IS VERY interesting for a customer of Heroku.
And if it isn't, than that customer is a fool for outsourcing so much of their system without even understanding the risks involved.
What would be interesting to me about an Amazon outage being behind a heroku outage would be to keep a tally, and if heroku didn't manage to build in more reliability to be resilient to even an amazon outage in a particular region, to question whether they were a good fit.
Please take this outage as proof that you need to build our your own infrastructure and hire your own operations team in multiple geographic locations.
In the mean time, we will continue to focus on building new features and products that our customers love on our EC2, Heroku, and cloud based system.
This is fiction.
Also: Most downtime is caused by bad code deployment, poorly-conceived network or system configuration changes, and sysadmins with fat fingers. Do you really think your hired talent is going to be better than Amazon's hired talent?
Honestly, it's not that hard to set up a server, and furthermore, it's just not that different to maintain a hard server than a virtual server.
And the most important part: when you have your own hardware, you at least maintain control over every aspect of your systems. The value of this cannot be overstated.
I'm guessing you haven't used Heroku. Server setup is "git push master".
you at least maintain control over every aspect of your systems
There are still plenty of things you don't control - network feeds to your cage, continuous power, bugs and failures in the hardware you buy. You cannot provision new systems without either buying machines or having a hot standby, and somebody needs to make a trip to the cage. If you're getting hardware by the month from a service, you will probably have faster turnaround (hours not days), but once again you're giving up some control.
I have an ancient quad-900Mhz Xeon with 6GB RAM (customer does not want to migrate) that has an uptime of 1600 days, and for which total network issues during that period was a few hours (wonky power to switch).
Cloud is too often comparable to "vapor" in terms of the claims of redundancy and availability.
They also try new things with large amounts of data. This requires scaling up to additional machines for hours or days at a time, and then scaling back some or most of them when they optimize new ideas and services for production.
When the new services are a hit with clients, traffic increases and the whole cycle starts again.
Now, what I think the original post there is speaking about is for mid to large size enterprise companies that have stable, but significant traffic. In cases like this they must do a cost-benefit calculation because you risk a lot if you don't. Then the cloud might very well not be the right solution, because costs could be 10x more than anything else... so the answer in my mind is not always clear.
Sure, EC2 is probably best for a startup that expects to double every week from a nontrivial starting point and has large machine resource needs per user (viral video startups, for example). The vast, vast majority of startups won't have anything that resembles that kind of growth graph, though, and thus shouldn't blindly follow what the Pinterests of the world do. It's a completely different type of demand. If they find out that they actually are going to have double digit daily organic growth percentages, then they can switch to EC2 before it gets out of hand, but otherwise, it's premature optimization.
Looks like they didn't actually do any of the remediation steps.
Somehow Heroku doesn't seem that great to me.
(And that comes with excellent support - for example, I asked if they were planning to offer Python and they said "Sure, just gives us a couple of days to set up a machine with Python for you", even when I was only interested in the cheapest plan).
I know they don't serve the same market, but I find it strange that a service that costs an order of magnitude more doesn't have a better uptime than cheap shared hosting.
No point bemoaning a lack of decoupling, if you don't actually use it.
There's a notification at the top of the page for me with that message, but it didn't appear in Chrome. Session collision maybe?