Thank You Heroku, or "How To Eliminate Sysadminning"
blog.dougpetkanics.com
blog.dougpetkanics.com
Those are two bold, scary statements.
Regarding the second though, I truly believe that the cost + time savings are tremendous from using their platform vs administering your own early on in a startup product lifecycle.
Except... reality is harsh and ugly.
You want the guy with the sysadmin foo on your team from as early as possible. Because if your thinking goes along the lines you just described then clearly you don't have him yet, nor someone who told you what slippery slope you're about to tie yourself to.
If you have the admin guy then he will find you a cost effective platform to start out with quite effortlessly. Which might be heroku in some cases, but is usually just a bunch of rented or virtual servers that he sets up over a weekend. In the latter case it might actually cost a few bucks more initially than the "heroku free plan" - but he'll explain to you in kind words why he thinks it's worth that in the midterm. And he'll probably be right.
If you don't have him, then heroku can't save you. Heroku will run your stuff for a while just until things get interesting and worthwhile load starts to build. That's the point where the sum of your mistakes brings it to its knees. Due to basic mistakes usually. Related to simple things like file descriptor limits, stuff about TCP connections, or a basic understanding of how disk i/o works and how to craft the SQL to make it not hurt so much.
At that point all you can hope for is that you're already profitable enough to afford not only that guy you skipped on initially, but rather his bigger, hairy brother, which is the only kind generally willing [and able] to take on "search & rescue" gigs. For an adequate sum.
At this point your friendly "pay-later" route has turned into a nasty "pay 10-20x now" roadblock. In an "Insert coin to continue" sort of way.
This is not about what heroku can or can not do. This is about what kind of skills you need on the team for a web-startup. Don't skip on the admin. Or rather, don't skip on at least one guy who has been doing the admin thing on a real live app for a bit and also during some not-so-good times.
What if you are not a making these mistakes, how is Heroku (or another service like it) going to bring you to a harsh and ugly reality?
It seems Heroku is fine if I don't want to monitor my own sever all the time and deal with the maintenance that goes along with that.
Well, if you are running a startup from scratch - you ARE making these "mistakes" because you certainly aren't focused on the minutia of performance tuning and significant scalability concerns.
Most startups at this point are still trying to get the aircraft off the ground - much less thinking about switching on the auto-pilot for cruise.
With that, this is really, really good advice.
In sysadmin-land, all my pain comes from software that was never written with management in mind. Assumptions like 'all TCP ports are open, all the time, between all servers', 'we can put things wherever we want in the filesystem', and 'the server should be configured just like my local workstation' make upgrades and deployment a nightmare, and are nominally difficult to change in a codebase with more than a few iterations underneath it.
Oh, and my other favorite pet peeves: Applications that provide no easy way to verify whether or not they are, in fact, up or down, and having debugging information dumped into the logs marked as errors, rather than as debug messages.
Keep in mind that these won't bite you in the ass initially, when you're only on a single server, or on a small, tightly controlled cluster of boxes. They'll torpedo you when you need to scale up, and you'll get a second shot in the boilers if you ever need meet any number of regulatory standards for various industries (PCI, HIPAA, etc.).
Fixing these early-on is easy, and you can enforce scaling-and-deployment friendly coding practices through automated unit testing and CI. Won't even really cost your coding team any extra time. Fixing them down the road, after you've got a few thousand live customers and a big pile of codebase to dig around in, is... difficult, at best.
One good thing about Heroku, though, is that you essentially have to do a lot of these things properly from the beginning because they enforce a shared nothing, tiered architecture. For instance, you can't write to the filesystem, you can't set up some poorly secured e-mail server, you can't have arbitrary daemons running. In general, every process that isn't core to the web stack must be executed in a completely different tier.
Obviously you can still shoot yourself in the foot with poorly written SQL and other things, but a lot of the "sys-admin" type mistakes are avoided by enforcing best practices when you start with Heroku.
I seem to mention this frequently, but I'll mention it again: there are many business models for which you can reach any level of "worthwhile" business results without ever taxing your setup. The whole notion of "Make a service on a shoestring, get popular, race against time to keep up with the volumes of traffic possibly crushing you before you can raise money to rearchitect the service" is the exception, not the law of nature.
This overrides the middle-click behavior of my browser to open the link in a new tab. I'm trying to open your outbound link in a new tab so I can continue to read your content AND later read the page being linked to; the analytics being surreptiously added to every link overrides and prevents this.
I have had poor results getting customers to use Heroku. I tried twice, and both times the extra cost over running your own EC2 instances convinced my customers to eventually spend much more money having me set up custom infrastructure - not a good decision unless you expect to have a very high volume site (I hope they are not reading this :-)
The nebulous cost of more developer and admin time is something they can carefully ignore in their cost calculations.
Bottom line is that people dream that their web app will attract millions of users, and they want to plan for outstanding success. While I am sure that Heroku must have customers with large user bases, the sweet spot seems to be for moderately sized web portals that you sometimes need to scale up on demand.
App Engine is the opposite: you're paying for the usage, so they spin up/down as much capacity as necessary to satisfy the volume you're willing to pay for.
If you are paying more for the dedicated cluster options, then I agree with you that they should not be spun down.
Certainly, it is OK for them to spin down the free option dynos if they don't get HTTP requests for some reasonable time period.
On the low-end I can get a VPS for $10-$20 month and get a Rails app up and running for a significant number of users with just that modest outlay. I can install SSL and any software I want and have a predictable amount of resources to play with. Yes, there is some sysadmining overhead, but setting up Nginx w/ passenger is like an hour of work once you've been through it a couple times, and similarly a lot of the add-ons that Heroku charges a monthly fee for are just a small one-time time investment. When I'm trying to bootstrap something really small the last thing I want is ramping up a significant recurring cash costs just to run some open source software that is really not hard enough to setup or maintain to justify recurring costs.
On the high-end when I'm using a lot of server resources, I'm paying a growing premium for a given amount of resources. Now don't get me wrong, it's nice to be able to magically adjust for traffic spikes, but if I have a consistently high amount of traffic, once again I'm paying a high recurring cost for actual resources that are a fraction of the cost, especially if I need any add-ons which are not inherently resource-intensive—maybe I'm actually just paying for SaaS of essentially open-source components. Actually I have this problem with EC2 and S3 in general to some extent, though much less so than with Heroku because the markup is lower and I've still got a large degree of control.
The benefit of Heroku is that they maintain an up-to-date and well-tuned Rails stack w/ add-ons on top of EC2. There is definitely value there. But in all cases there is this downside of overhead and flexibility. The clincher for me is that ultimately Heroku may end up not supporting something I need that would be trivial open-source stuff on any UNIX VPS. At that point I will need to migrate off of Heroku and all those supposed sysadmin costs hit me full in the face all at once rather than being amortized over the life of the project. Maybe I'm just turning into a cranky old man (at 31!), but VPS or dedicated servers still seem like a better value considering all risk factors. If I had VC money and was trying to ramp something up really fast I might reconsider.
So I moved to Heroku to focus purely on the app. It's wrong to say they change for "every little thing"... most addons are free. But yes, hourly cron, background workers, unlimited bundles etc are charged for.
I'll always recommend to anyone that they check out Heroku because I think it's a real time saver if your app fits within their constraints. During development you'll usually save money and then the costs will scale (hopefully) with your income after launch
I'm sure Heoku is a great deal if you don't have the sysadmin skills. My skills are just good enough and I manage my own box. I bought it 5 years ago..4 cores, 8 GBs, 73GB RAID-1...for $5000. I pay $105/month to host it.
This is at least the equivalent in horsepower of what you get from Heroku for $500+ a month (that's being generous). So in a few years, I'm ahead of the curve: $11,300 vs $30,000. Add in the fact that I paid down $5000 in advance as a capital investment and I'm still ahead. You could do this today for under $3000 capital investment and have an even more capable box.
Sure there's risk I'll have a hardware failure (I did have 3 years 24 hours on site warranty, but that's still potential downtime and headaches). But there is also risk that over the course of that time you will want to run things on your server you did not anticipate when you started. It certainly was the case for me. e.g. I started my app in rails then switched to merb. Now I run a second site on the same server with mongodb and solr.
Bottom line, if you have moderate linux skills and know your going to be in "some business" that requires you have a server, buy your own. If not, Heorku seems a fine choice for some.
Can you recommend the company that you colocate/lease-to-own with?
m5hosting is bigger than when I started with them but they are still a "medium-sized" company. They seem to be solid and whenever I have a support request they are on it quickly...about half the time I get a response from the owner.
I generally follow the rule of picking business partners that are "right-sized" for me as a customer. For many services I don't want a company so big that I'm not worth their time after the sale...especially when it comes to server support.
I have never been to the data center and never seen my own server. I bought it on ebay and had it shipped to them ;).
I think that it would be sensible for a VC founded startup to eliminate risk and go for scale on day one with an internal infrastructure team + self managed, colocated physical servers.
Guys, let X/N people do the design, while (N-X)/N go around looking at your competitors' websites for "inspiration". What you wanna do is copy business plans, target markets and monetiziation strategies. Not goddamn pixel and element positions.
The first is only hourly cron tasks which means doing sweeping every 15 minutes is not possible. Not terrible, but sometimes a showstopper if you're trying to do background tasks frequently.
The second is that it's a read-only filesystem which means you either don't use the filesystem for your work, or you use S3. S3 is great, but is another service you need to monitor and pay for. You also will need to employ other solutions to compress CSS, JS files, or to create a cache.
Third, there's no shell. There's a console, like the Rails console, but you're not going to be setting things up yourself.
Heroku is a heck of a platform. I run tons of stuff there, even on the free plans with no problems. However, the limitations can be showstoppers for you depending on 1. how much hackery you want to do and 2. if it's even possible.
I find that the read only FS is a good thing, decoupling storage from app servers really helps wrt to scaling.
The real problem with hosting Spree on Heroku is having to spend $100/m for an IP for SSL. The actual cost for setting up ELB instances to get extra IPs for SSL is $20/m. That's a problem when you want to set up a bunch of differently branded storefronts. It's the main reason why I'm doing DIY with http://github.com/wr0ngway/rubber/ instead of using Heroku for the spree project I'm working on right now.
SSL alone is as expensive as a 1.5 GB slice from slicehost. I know it's not really their fault, but they need to work with amazon to get several IPs for each of their images, or something.
For instance, what about Heroku/EC2 prohibits me from using Authorize.net as my payment processor, provided I never store card data?
EC2 is bare Linux; you install apache/passenger/REE/memcached/postgres/etc/etc/etc. If you want to scale past one server, you're on your own there. You take care of backups, security, system monitoring, etc.
Heroku is a platform that abstracts you from these things. You just push out your app, tweak a dial or two on your Settings page, and they take care of everything else for you.
If you want CPU-by-the-hour, go with EC2. If you want to host a Rails app, go with Heroku.
My current focus is to look at the plethora of excellent SaaS applications that help provide the whole Rails ecosystem, then turn around and do the same for Django, with an eye towards expanding to handle other Python frameworks, and then into the Java space.
It did deploy on commit/push, bundled eggs, worked with mysql or sqlite and redis. Dunno if that sounds interesting to others though.
Basically, nginx fronts a bunch of Varnish caches, which are then in front of a "routing mesh", which queues up requests for individual "dynos". A dyno is a single-threaded instance of your application. When one dies or is migrated, another is deployed in its place. Multiple dynos are automatically spread across multiple machines. It's a very neat solution to automatically scaling hundreds or thousands of apps.
In other words, bundling the apps and pushing them out is just the beginning of a pretty fascinating architecture.
Talking to some other Pythonistas, there didn't seem to be much desire for a Heroku-equivalent system though - most ppl enjoy rolling their own deployments.
Going down to the metal may help squeeze the last 10-20% out of your hardware, but the the really interesting and challenging work in coming up with a scalable hosting platform is elsewhere: security, monitoring, process spawning/reaping, deployment, et. al. If working in Python gives you a time-to-market advantage, then go for it. You can hire a C hacker when you have enough business to make the improvement in your hardware utilization efficiency pay off.
Other than that? Lots and lots of python, of course ;) I'm only one guy, and there are three or four big moving parts that need to be written. Something to build eggs and push them, something to monitor and manage processes on individual servers, something to dispatch requests to the appropriate servers, and something to monitor and manage the database servers.
Even pulling back and looking at a minimum viable product, I'm still doing most of that work, just far simpler versions. Succinctly, it's too big for me to tackle on my own in the next six months.
The app-engine-patch team has essentially forked Django and made it work on top of GAE models. However, there are some disadvantages, including a huge perf hit on cold requests. I tend to avoid AEP.
I'm currently of the opinion that trying to get "all of Django" running on GAE is square-peg/round-hole. You still get a lot of mileage from the bits of Django you _can_ use on GAE, and it's modular enough that it is possible to build new GAE-specific session, etc. objects. I have several OSS repositories on GitHub that could serve as a decent Django + GAE template if you're interested.
In particular, I think it's an issue that there's no decrease in cost for dynos (though I be for extremely large sites there's special pricing). Because of this and because you get a freebie, dyno pricing is actually a little progressive (ie. your average cost/dyno increases as you get more dynos).
If you have 2 dynos, you're paying for one at ~$36/month, so your average cost per dyno is $18. If you have 11 dynos, you're paying for 10 for a total of $360/month, which is an average of $32.73/dyno-month.
But one thing that is making me think about VPS instead of Heroku is the ability to run multiple apps from one VPS unless I really need the power.
Does Koi allow you to run multiple Rails apps/websites per Koi, or is it 1 per Koi?
In any case, pricing is separate for each Ruby app you wish to run. You could run one under Blossom and the other under Koi, no problem.
"upload and go" is a convenience feature of Heroku, not it's reason for existence. It does make it easier to deploy Ruby apps but, again, it's not the reason it exists.