Ask HN: What's Your Worst Bug?
Some memorable ones for me:
1. Around 2006: hand-coded replication link between primary and secondary of an HFT component, each with its own data store; primary flushes its queue when secondary acks up to the message sequence number it has done a write() to disk; primary crashes in production, secondary takes over, rebuilds state from its own data store, but one sequenced message is missing; data stream cannot continue; pandemonium; and that is how I met disk write buffers in Linux. Write doesn't mean write until flushed to disk.
2. In the aftermath of (1), to solve the above problem, called fdatasync() after each message write by the secondary; production comes to a screeching halt; latencies go from around 5ms to 100ms; pandemonium; and that is how I learned the cost of blocking I/O in a single threaded environment. I ended up calling fdatasync() only once in every 100 writes.