It's been a while since I've used it at scale, but at my previous job half of our incidents were caused by Postgres randomly deciding to change the query plan for a call that was working just fine. Then I'd go in there and re-write the SQL a few times until it figured out what to do, rinse and repeat every few months.