The Switch from Heroku to Hardware
justcramer.com
justcramer.com
I'm sorry, but no.
You don't pay an ops guy or Heroku buckets of $$ for when things are going well, just as you don't pay $$ for software that only handles the happy case.
You pay $$ for someone who has fixed shit that went horribly wrong and has the scars to prove it. "That deep purple welt on my lower ego ... is where I only had a backup script and never tested the backups. This interesting zigzag is where I learnt about things that can wrong with heartbeat protocols ..."
Edit: though see below for a more nuanced discussion of reasons from OP.
Nope.
I don't need to suffer the same acts of terror operations people have gone through to be able to avoid, prevent, or recover from them. I'm paying myself to be the operations guy, as well as the engineer.
You're betting foregone engineering time against Heroku's value-as-delivered.
I don't think that the hours you will inevitably spend fixing the ops infrastructure will turn out to be very profitable.
For a large business, moving off Heroku to inhouse operations is probably justifiable, because they can capture sufficient value from a smart ops team to offset the potentially very high monthly bill Heroku would levy for their business.
But for a smaller firm, Heroku so abundantly oversupplies value compared to the bill that you will simply never be able to capture that value from internal work at anything like the same price.
It's like saying "I should stop buying food from the supermarket, it would be cheaper to grow my own vegetables!"
In a naive dollar-cost analysis this is true. In an actual consideration of whether raising a vege patch, bearing the risks of starvation, doing it poorly because you're new at it, spending the hundreds of hours of labour it will require -- means that letting a professional farmer do it is much smarter.
Gains from trade are gains.
I have a pretty strong opinion on this given how much time I've already sunk on similar work: http://chester.id.au/2012/06/27/a-not-sobrief-aside-on-reign...
And for my secondary startup I am thinking I will just let Heroku handle it.
https://github.com/heroku/wal-e
It is not the best program I ever principally authored (since I know you are very discriminating Python programmer), however, it does work very well and has an excellent reliability record so far, seemingly for both Heroku and users of wal-e per most reports. I also tried quite hard -- and in this I'm pretty satisfied given both the feedback I've gotten and my personal experiences -- to make the program easy to use and administer, for one server and operator to many servers and one operator.
If someone is not doing continuous archiving and can't use our service for any reason, we do try to urge them to use something like wal-e or any of the other continuous archiving options. And to that end, I tried to make a pretty credible wal-e set-up only take a few minutes for someone who already knows how to install Python programs (I would love for someone to contribute more credible packaging).
To be in control of your venture means knowing all the corner cases.
I'd trust dedicated hardware more than Heroku as when it does go wrong, you're at the mercy of yourself and not others (other than the colo facility).
There's nothing in this justification that doesn't also apply to Heroku. The cost structures just aren't significantly different at a two machine scale. However, as people keep pointing out, the roll-your-own-cloud approach requires that you build and maintain a bunch of infrastructure that Heroku has already built for you, or you forego redundancy and fault tolerance that Heroku has already built for you.
The best lines of code are the ones you don't have to write.
There's an I/O penalty for working on AWS, but it's on the order of tens of percent, not hundreds. I suspect that your original problems were related to working set size relative to cache (since Ronin => Fugu bumps cache by over 2GB, and you said that Fugu was working well).
Heroku's largest database has a 68GB cache at an (admittedly expensive) $6,400 a month. But even so, $6,400 is a small expense for a growing web application. A mediocre developer costs more than that. Trading off server cost for developer cost is an asymptotically bad bet.
The database definitely wasnt unreasonable at $400, but for a bootstrapped project (especially something thats a side project for me), that was a big consideration.
I probably would have toughed it out with Heroku if I could have gotten things to perform better. At one point I was running 20 dynos trying to get enough CPU for worker tasks to actually keep up, and I unfortunately couldn't solve the bottleneck to where the cost was reasonable.
The application isn't typical (what Sentry does pushes some boundaries of SQL storage for starters), but it was costing too much of my time to struggle with optimizing something that really shouldn't have needed that much effort.
I definitely like the redundancy provided, and the ability to add application servers with zero-thought is a huge plus, I just couldn't justify the cost of the service in addition to my frustrations/time of trying to scale it on Dynos.
Something doesn't make sense. If you "addressed" your I/O problem, your CPUs were therefore all busy doing something much, much slower than a disk read/write, in software (which would have to be both obvious and unbelievably horrible). If that's true, something pathological was going on in your code. I'm going to assume that you would have noticed it -- swapping, for example.
So let's go back to I/O: if your database was slow, you might observe something superficially similar to what you've described: throwing lots of extra CPUs at the problem would result in lots of blocked request threads, and appear that your dynos were all pegged. The exact symptoms would depend on your database connection code, and your monitoring tools. But in no case would throwing more dynos at a slow database make sense, so I'm going to assume that you didn't do that on purpose (right?)
Given the above, I still can't meet you at the conclusion that abandoning Heroku was the magic bullet for your problems. There's not enough information, and it doesn't add up. My money is on one or more of the following: DB cache misses (i.e. not enough cache); a heavy DB write load; frequent, small writes to an indexed table; or pathological memory usage on your web nodes. And if it turns out that the cause is due to I/O, you've only bought yourself a temporary respite by moving off Heroku. Eventually, you'll get big enough that the problem will re-emerge, even though your homebuilt servers are 10% faster (or whatever).
EDIT: Aha! Your comment in another thread actually explains your problem: you were swapping your web nodes by using more than 500MB RAM (http://news.ycombinator.com/item?id=4458657).
More importantly, this a few months ago I made the switch, and I don't remember the specifics of the order of events. I can assure you though that I know a little something about a little something, and I wasnt imagining problems.
(Replied to the wrong post originally, I fail at HN)
Since you've made it clear in another thread that you were actually running out of RAM on your dynos, I imagine you were running into trouble. There's no need to be snide about it.
Bottom line: you hit an arbitrary limit in the platform. If heroku had high-memory dynos, the calculus would be different. In the future, instead of arguing that your homebrew system is better than "the cloud", you could just present the actual justification for your choice.
That's rather optimistic.
The EC2 ephemeral disks normally clock in at 6-7ms latency, that's >13x slower than dedicated disks.
EBS clocks in at 70-200ms latency, that's >5x slower than a dedicated SAN.
And that is under optimal conditions. In reality the I/O performance on EC2 frequently degrades by orders of magnitude for long periods of time.
Um. Are you comparing hard drives to SSDs? Rotational latency for a 15k drive is a couple of milliseconds. Seek time for server drives varies from 3-10ms.
EC2 disks are slow, but there's no way they're 13 fold slower than your average server drives. And 6-7ms is just about on par with commodity hardware.
No, but I mistyped the numbers. That was supposed to read: 60-70ms.
Heroku at least uses EC2 which will be far more reliable over time.
It was the fact that Heroku offer an environment to run fairly standard Rails + Postgres that made me pick them over the more unique and harder to move from Google platform. Even though I was starting from scratch.
It's always good to have an exit route.
Those of you running startups that don't colocate your own hardware, and don't run in a cloud, where do you rent servers from these days?
Most of my stuff is at Softlayer, but their RAM pricing is killer ($25/mo/GB).
For Sentry I'm using Incero, but will likely be switching to Hivelocity (or something similar), as I currently don't have internal network and that's a big, annoying deal for me.
Also, if you're loooking for deals or more information on hosts: webhostingtalk.com
SoftLayer have much better support, geographically diverse locations and a wicked suite of offerings. They're also giving startups $1000/month credit for 12 months, I got almost instantly accepted... I was still asking questions and then the rep was all "I made you an account here's the details".
I used to have a really good deal with SL about 2-3 years ago, and when I was looking around for Sentry they were my first stop. Unfortunately the prices had gone up a lot while I was gone, and for a reasonable DB server it was looking to be pretty pricey.
As for desktop-grade hardware - I haven't seen this mentioned by anyone else. Any links to back it up?
If that's true (and I doubt it is), it's avoidable if you use their extremely well priced colo option (which is what I'm going for).
Yes, "high latency" was referring to the latency. A CDN for static assets doesn't mitigate the fact that the initial request, and everything dynamic is on the wrong side of the planet for most startups' users.
> Any links to back it up?
Hetzner.de. They list the hardware in the boxes. It's desktop processors, desktop motherboards, desktop hard drives and non-ECC RAM in all the cheap server lines.
http://www.hetzner.de/en/hosting/produktmatrix/rootserver-pr...
7200RPM hard drives, Core/Athlon processors and non-ECC RAM don't belong in servers.
http://arstechnica.com/business/2009/10/dram-study-turns-ass...
According to Google's study, you would expect about 2 memory errors per day per server running 24/7. You need ECC RAM.
> I doubt it is
It's rude to publicly call someone a liar without evidence.
Re: hardware. The choice is nice to have and their EX6 and up packages are "proper" server grade (as in Xeon and ECC). EX6 machines start at EUR69 which is fantastic if you ask me.
Edit: as for rudeness - chalk it up more to a disagreement over what constitutes "desktop-grade hardware" (and your omission of their higher level hardware offerings)
After a bit more chatting with a sales guy I managed to get the price per gig of RAM well below that.
Softlayer is a bit expensive out of the box, but I've found them very open to negotiation.
Now, from UP2VPS[1] you can get a pretty good deal (1GB memory, 1TB transfer) for just $6/month.
[1]: http://up2vps.com/
But if you don't need a support, their service is very affordable and good.
So, yeah, ops isn't hard at all, if you don't fucking take the time to do it right.
You can even add to /etc/hosts and the computers using it as their DNS will resolve it. Depending on how much control you have, DNSMasq will also function as a DHCP server and TFTP server from which you can netboot other servers and do such nifty thing as automatic reinstalls. Useful if you have a separate, internal network and want to set internal IPs, too.
Even with a local DNS server, there has to be some overhead though.. OTOH, avoid premature optimization etc..
Ps: to the op, not parent.
That DNS mapped to one node, which will rarely change. If it does, I can spend 30 seconds and deploy a new config, rather than worry about using virtual IPs or anything more complex, let alone with having to configure a cache which would have the exact same problems (delay to change).
Guess what, it's 2012 and the same shit that worked back then works just fine now.
There's a fine line between doing things right, and doing things just to do them.
I suppose my real objection is in not Doing The Right Thing. Yes, it may work for your current two box static setup. But small choices contribute a lot of debt that you or someone else has to pay off later. Would be a shame if someone took the message that DNS was just "overhead".
Is this not similar in strategy to choosing hardware and spinning up your own stack?
EDIT: I getting at the rather sweeping statement against all cloud providers based on a specific Heroku problem