We need to write up the incident report; it’s not clear whether the root cause was a bug in postgres (potentially years ago, and only caused problems when we started using the corrupt piece of the index by coincidence) or due to potential corruption from a HW failure years ago.
The rooms affected should now work again after reconstructing the lost data. No data should have been lost over federation though; the room history should have synced up as normal. Also, only room state (not msgs) was effected. Do you have a bug report anywhere for the missing msgs?