Parts of Heroku are down
status.heroku.com
status.heroku.com
Heroku's provisioning, build, and remote console services have been down for a few hours. Luckily my production app is still chugging along pretty well. There was a time this afternoon when I wanted to restart one of my six dynos (because it was half-dead) but was unable to due to the API lockdown. That meant every sixth request was sent to an app server that was likely to barf on it for a while there. The dyno got better a few minutes later though.
Postmortem edit:
API outage is over after a period of about three hours (that I was aware of anyway). During this window my app's response time was about 1500ms. Now that it's over it's down to about 250ms, which is on the high side of our daily variance between 150 and 250.
Unfortunately for us, we pushed out a marketing newsletter right about the time they shut down the API. Looks like we're still getting decent sales though!
I've worked and built enormous projects in the past and hosted them myself (on hardware, and on providers directly (Rackspace, AWS)), but have always had more headaches, wasted time, and downtime when doing things myself (and with a team) then when I'm using Heroku.
Regardless of the occasional incident Heroku has, I'm still 100% a loving user. The people over there work super hard on tons of stuff, and do a great job at keeping millions of applications up.
Keep at it! <3
Having been a heroku user for quite a while in the past I'm not surprised, these sorts of issues are sadly common with heroku.
Heroku also has no transparency into its architecture, which is why no site on Heroku will be able to be PCI compliant after Jan 2015.
I'm not optimistic about Heroku's ability to continue to be a leading PaaS.
do you have a reference for that?
Heroku has never been able to pass a PCI Level 1 or Level 2 audit, but as of Jan 2014 it will also no longer pass level 3 or level 4 SAQs.
Hit and run comments aren't really useful. Some detail and if possible some suggestions for improvement really help. Even if they can't be implemented for some reason.
It's mostly documentation. If Heroku has built a secure system and documented it adequately, then it would easily pass SAQ A-EP.
One easy example: Maybe Heroku keeps logs of all HTTP requests that include params containing credit card numbers. Nobody knows.
Kind of surprised Heroku went down for so long in the middle of the day. Can't imagine any serious large scale services wanting to stay on a platform with that availability for very long. We're not large enough yet to move, but I think AWS is going to have to happen at some point.
I'd echo rdegges point from eariler, Heroku has been (and will be for the foreseeable future) equivalent to a FTE in giving us the ability to build things and not worry (too much) about the hosting.
Edit: note the response time increase _during_ the downtime event.
Builds are back up for me after a few hours of downtime.
Of course, the IT team would probably hate being held visibly accountable like that. But hey, I guess that's why they are working in corporate IT as opposed to a well-known and successful cloud service provider.
Is your IT team compensated at the same level as well-known and successful cloud service providers? If not, why would you hold them to the same standard?
In general, "Foo is Down" doesn't make for very good HN posts, because while it matters to users of Foo, the fact that Foo is down usually isn't intellectually interesting (which is what HN is looking for in a submission). Postmortems about why Foo went down, on the other hand, are often fascinating. Conclusion: we should usually wait for the postmortem.
edit to add: Thanks for the transparency in moderation! I had thought this thread was poorly titled too.