It's entirely my fault. I screwed up in a few ways, which is what you always see in something like this.
Im on an overnight business trip from Boston to Dallas.. My Real Job wanted me to tour a new datacenter down here.
In any event, I've been running straight, without sleep since yesterday AM when I flew out on a Red Eye, so I.. I'll admit it.. I went to bed.
I was sleeping peacefully from 1AM to 4AM, cradling my old-school, screaming-loud pager secure in the knowledge that all my little servers would alert me if they had problems.. I have Pingdom, PagerDuty, Nagios, Munin, I've got monitoring 16 different ways..
When I woke up at 4AM (3 hours is enough. Why not?) to make my flight back to Boston, I see I've received several hundred text messages. Hrmm.. What is this?
I double-checked the work-servers, but those were fine, and have Other staff to help watch them, and escalate. Checking email, I see Pingdom is patiently explaining that RoboHash is down.
"But WAIT?!", I can hear your exclaim! Why didn't your pager wake you up?
Well, it turns out that old-school pagers are.. Regional. Once I'm outside of NewEngland, it's a cool looking retro piece of trash.
The text messages arrived, but my sleeping brain was able to peacefully ignore them in bliss.
As to what actually happened? Linode migrated my machine to a new datacenter, and rebooted it. I had a bug in my init script, so it didn't start up properly when rebooted.
Basically, I need to create a Ramdisk, then copy the code from the stable position on HDD. This is because a bazillion hits/second would overwhelm the disk if I loaded it manually from disk each time.
Anyway, Nginx was started (hence the error), but the init script was trying to start up my python code, without creating the ramdisk first, so.. RoboFail.
Sorry about the problem. Imagine my embarrassment.
On the plus side, I'm now in a TSA-approved boarding area, so I have a good 20 minutes to fix the init script before my flight back. ;)