I left Heroku for OpsWorks and shaved 40% off response times
stefanwrobel.com
stefanwrobel.com
A friend of mine and I have been working on a startup for the last 8 months. It has been tough as we have been working full time for a company and hacking like crazy once we got home, or on weekends.
We chose to use Heroku because of the dead simple workflow that would allow us to spend time developing features instead of maintaining the infrastructure. This may sound like a cliché, but boy was it a good decision. The platform is rock solid, deployments are a bless and the add-on system is just one of the best time savers you can get.
We know OpsWorks (or earlier Scalarium) from our full-time job, and we know how cool it is. But we also know that it requires a lot of time and additional effort to keep things running.
This may pay off if you run a company which can afford it, but you should think about it twice if you have limited resources (and other, more important things to do!).
Even now, after we have launched, we still see no reason why we should move from Heroku to OpsWorks (or another platform). We have made some optimizations and clever (to us) decisions to avoid running into scalability problems. Maybe we will think about migrating once we get a ton more clients and even more traffic. It is cheaper to just add some dynos than spending time on server administration.
In the end, it's all about making the right decision, so don't hurry and if you think Heroku is too expensive for your business, maybe you are not making enough revenue from your customers. On the other hand, if you have the money and time, than go for it.
Concurrency is only half the battle with tail latencies. Of course dedicated CPU helps as well :)
Heroku is offering 2X and soon, 4X dynos; did he experiment with these?
The author also sounds rather excited about nginx serving static assets. The Rack::Cache problem on Heroku is easily solved with a CDN layer (and Heroku even recommends you do so.) (Cult Cosmetics /is/ set up with CloudFront, so the author obviously knows this.)
There are a lot of Rails architecture smells in this article, such as worrying about the 60-second boot timeout and the 30-second request timeout. [1] I have to wonder if the time invested in learning Chef and managing the architecture may have been spent making the app more performant.
Developer time is really, really expensive. Not only can you bill $100-$150 hour instead of learning how to use Chef, you could be shipping features for your own products, which can bring you passive income forever. A lot of companies are successfully running on Heroku, and it seems strange to throw everything out for the explained reasons.
[1] Specifically, adding 'require: false' to your Gemfile and only loading heavy libraries when you need them greatly helps with boot time. Ideally, you won't run any image processing libraries (like RMagick) except in background jobs, and your main app should not even need these libraries available. Same goes for PDF generation, scraping technology, or any other heavy processing.
The 30-second timeouts either need to be handled via background jobs (the author mentioned they use sidekiq) or uploading directly to S3, side-stepping the web process entirely.
Also, we're an ecommerce site so we rely on a lot of external services (specifically for payments) where we have to run the call in the web process ... hence the issue with 30s timeouts.
I know you probably have good reason for this, but this smells like an architecture problem. Whenever you are talking to an external service it is a good idea to have your code laid out such that you can do that in the background. I understand that if you didn't plan enough architecture and maybe even product design to make that feasible from the get-go it is an expensive thing to build, but regardless of who is your hosting provider sooner or later you'll likely need to do it.
One thing I've always liked about heroku, at least in concept, is that it forces you to have good architecture. Some times it is a bit of a pain, but I think this pays off quicker than you seem to realize.
Generally, you use JavaScript to actually tokenize the card. You then save that token in your actual application, and then use it to actually charge against.
This model works just fine with background jobs.
Looking at the c3.8xlarge instances, they're specced pretty much exactly like the servers I just built for $9k/each, except they have half the ram (60 vs. 128) and half the storage (2x320 vs 4x320). And they cost the same to rent for 6 months as it does to buy one, including coloing.
So, if you need more than a fraction of a physical server, EC2 still loses pretty hard on the pricing. It seems as though it has become competitive with Softlayer, though, at least.
EDIT: Ah, but if you pay $11.6k upfront to reserve it for three years, it reduces the monthly bill from $1728 to ~$370, which is a much better deal than Softlayer offers. That's actually a pretty reasonable premium compared to racking them yourself, though you lose flexibility and the hardware is worse than what you'd get yourself. Well played, Amazon.
FYI, you've been hellbanned, Sssnake.
They use slightly better processors for the C3s than I use (2x8 core 2.8 vs. 2x8 core 2.6). I'd assume that if you get the largest instance, you probably have the whole machine to yourself and don't have issues with bad neighbors, so it's probably actually slightly faster than mine.
That being said, opsworks is still based on chef, which is still cumbersome. For most people Ansible is far easier to deal with if they don't fit into the opsworks box. The overhead of starting from scratch with ansible can be made up for with the faster overall development time.
EDIT: ohai Andrew! Thanks for the comment.
If you are about to embark on the journey of provisioning for the first time, or if you are looking for something a little more straightforward than chef/puppet (in my opinion at least...), I highly recommend checking out Ansible!
"AWS OpsWorks can scale your application using automatic load-based or time-based scaling and maintain the health of your application by detecting failed instances and replacing them. You have full control of deployments and automation of each component."
How would you suggest using Ansible to cover that side of things?
I returned a couple of days ago to look at the current state of things and I have to admit that I was impressed. Technically I really liked ansible for quite some while. Now it even comes with decent documentation. I'd back the suggestion to check it out.
Maybe that's not a problem since I rarely found an open source ansible playbook for anything I needed and had to always write my own.
One downside is that I can see how tracking down role dependencies would be tricky if you had many different roles. I also haven't tried at scale.
re:chef
I don't think this is necessarily a fair assessment and true for all. There are a ton of community cookbooks out there, and a very large community now online to help people with Chef. I've been using Chef for years and like anything else once you get over the learning period, it becomes second nature. Any tool that you use will have some learning curve to it, and so the trade offs that make it cumbersome for some, make it easy for others.
Also, you could technically Chef in Ansible in OpsWorks(if OpsWorks met other needs for you), but thats just crazy talk haha
All in all, I can say that I'm happy with the transition, but I'm also running one Dyno which I'm far from maxing out. I feel like for a larger product, Heroku would be more problems down the line than it's worth.
By the way, great tool, thanks for building it!
It's easy enough to understand how this happened as OpsWorks came about via an acquisition, but consistency would certainly be nice.
Presenting overlapping yet incompatible features under the blanket of AWS is confusing and potentially pretty frustrating.
This is nothing on AWS OpsWorks, as they greatly simplified interface into managing clusters (for deployments, organization).
I found about this recently as I updated a community cookbook(python), as the Chef community is moving forward (with Chef 11+) that 1.8 has reach EOL and 1.9 is stable (2015) and 2.0 current, and 2.1 is new-hot-ness.
Did the author investigate why the response time were that slow to begin with?