> If there was any manipulation of community notes in the last 2 days, how would we know?
You can't know until the data is published. 2 days isn't that long though. Just wait a couple more days for the next data dump, then run the algorithm and compare the results to what the X UI was showing at that time.
> If there’s manipulation of this data before it is published, such as ratings or notes never hitting these data files, how would we know?
That would be a bit more sneaky than just outright removing notes. As you noted, you'd need a user whose ratings or notes were omitted from the dump to notice and come forward. Or perhaps with careful analysis you could prove that the manipulated data could not have resulted in the allegedly removed note being shown and then later not shown, indicating something fishy happened.
Theoretically if X wanted to improve on this system, they could go even further and implement something like certificate transparency (append-only log verified by a publicly distributed merkle tree), or create an independent third party organization that users interact with to submit and rate notes, rather than that happening through X's UI. Given the threat model though, I feel like the UX and complexity trade-offs of that wouldn't be worth it. Open sourcing the data and algorithm as X has is already far more transparency than we get from any competing social media company.