Bitbucket Downtime Post-Mortem
blog.bitbucket.org
blog.bitbucket.org
Cache link: http://webcache.googleusercontent.com/search?q=cache:http://...
Disclaimer: haven't read and can't read the article. EDIT cache is here: http://webcache.googleusercontent.com/search?q=cache:http://...
(Totally OT: would it be correct to say "haven't and can't read"?)
Aesthetically it doesn't scan well - I think adding it in parens would be the way to go if you want to make that kind of emphasis.
Haven't read (and can't) ...
If the question is about whether it is grammatically correct then I wouldn't like to say - I never really paid much attention to those rules. :)
We recently migrated our WordPress instance to a new machine, but not all of the configuration was moved over. We've put caching back in place, so it should be back to normal.
http://rsyslog.com/doc/queues.html
http://rsyslog.com/doc/rsyslog_reliable_forwarding.html
(This isn't, however, to say the documentation on this stuff is easy to find / understand. The rsyslog documentation is definitely not the best.)
I disagree with the conversion to UDP, and consider it an anti-pattern. In general, whenever you have many front-ends sharing a back-end service (in this case syslog), some flow-control is necessary - and UDP with no flow control will one day come back to bite you.
Consider an attack or peak load event on the front-ends, now in addition to the direct capacity problems that that may induce, you also impact your backend and control-planes by flooding a network or service with uncontrollable UDP.
As others have noted; it would be better to use non-blocking TCP I/O with some bounded queuing for retries. It's also generally a good idea to use a LIFO queue, so that when the backend is restored after an outage, the recent data takes priority over old data (for logging information the health of the live system is most important).