Dribbble offline due to Postgres problem
twitter.com
twitter.com
In such a case, the very first thing you do — aside from verifying that your backups, replication, and WAL archiving are all working — before you even try diagnosing your problem any further than "something weird is happening" is make a filesystem-level copy of your PostgreSQL cluster. (If you're running LVM, ZFS, a SAN, or any other thing that lets you take an instantaneous filesystem snapshot, so much the better. Do that, and then copy it.) Then, and only then, should you even contemplate trying to un-fuck your database.
I cannot stress this point enough. That means make a tarball, or cp the directory, or rsync it, or whatever disk file level tool you prefer, and specifically and emphatically not pg_dump. (It's probably not going to make things any worse, but if you do have corrupted disk pages, pg_dump isn't likely work, anyway.)
Flailing around trying to fix things can sometimes make them worse. If you're working on your already broken data, and break it further without the safety net of a fs-level backup, you're ... well, you're worse off than you were five minutes prior, aren't you?
Still true today and as I say backup don't fuckup and plan for the worst and let Murphies law work in your favour.
I don't understand how a "real" company, even in todays overheated environment of soon-to-fail hipster startups, doesn't do (at least) one of these two things: a) have competent employees on staff that are intimately familiar with their critical infrastructure and how to support it or b) pay, yes, gasp, pay some other company for professional support services. A glance at the Wikipedia entry for Postgres shows some possible companies that do that. Or, as already mentioned, there are mailing lists.
But twitter? Really? That's support? For someone other than a hobbyist running a website in their spare time?
Go ahead, flame me. But my first impression is "amateur hour". My apologies if dribbble.com really is an amateur effort.
Just google "reddit status" and check the first result.
No it isn't. It is akin to shouting at random people on the street "help I don't know what to do!". Asking on IRC on a mailing list is going to a group of people who are there specifically to help people with problems on that subject.
These let you go back to any point in time. If you ran 'delete from orders where id=id', you can restore to the transaction before you ran that command.
http://www.packtpub.com/how-to-postgresql-backup-and-restore... contains more information.
Also, postgresql 9.3 (out in a few months) supports disk page checksums which can detect filesystem failures immediately.
If you are hosting your database on ec2, using only wal-e means that all of your data is hosted on amazon. If they were to cancel your service, you'd lose your data.
Running pg_basebackup+pg_receivexlog on a different provider is cheap insurance against that.
Be sure to test how fast wal-e can restore your data, btw. Restoring from s3 was significantly slower than restoring from a local disk (in my testing a while ago).
What I've found is that most of the files can be grabbed really quickly, in the 3mb/sec range, but there's always a handful that run at 300k/sec. Running the downloads 8 or more at a time tends to help with that so that we're mostly maxing out the local network.
My guess is hard drive failure. Hope they have good backups.
If you have a real business you need a multi-generational backup scheme of some kind.
Often, the need for a backup arises from the need to go "back in time". RAID and replication offer no solution for this. Ok, so depending upon your replication setup, it may offer some help, but the steps required to go back in time usually involve starting from a recent backup, and using files involved in the replication to "replay" recent changes to get you back to the moment in the past you wish to return to.
mysql> drop database production;
No way raid/replication will save you then. Same thing for some rogue delete statement gets in the code that deletes too much data.
Raid isn't a good backup because groups of drives are known to die at once or around the same time.
In normal life, your secondaries just chug along, consuming WAL logs.
Should you do something like drop/truncate a table, you can start with your base backup and replay the logs till just before your fat finger.
Still least I hope they didn't have transaction logging onto the same discs, seen that in horror because somebody had large raid array and did not think they needed the expense of another.
Postgres community is quite responsive and usually give very effective advice, as long as you use the proper channel to communicate.
In the interest of saving time: browser / version / OS? Any browser extensions installed?
For clarity: Mac and Windows Chrome sans extensions, but these days I'm running a few like Adblock and Chime.