With a few sources you can build an accurate single model of an event. You then edit that event as you see fit and then generate as many "independent" sources from different viewing locations as you want. Then upload them to different social media sites at different times.
And the real verification would come from ultra-high-res captures from drone cams, self-driving car cams, traffic cams, etc that are constantly running and recording everything, so that you can't realistically fake soemthing without simultaneously compromising uber, amazon, walmart, etc.
This only works in large events. What if you film a politician lying at a small event, but his image team makes 20 fake videos showing the event in a different light? I guess you're the liar then.
Secondly, your analogy is wrong. The new hypothetical you present is equivalent to the old model where you are all giving verbal accounts and his staffers lied. If you are going to compare to an old-tech scenario where you one-up them with a more detailed recording than they have, you need to do the same with new technology to keep the comparison accurate.
Thirdly, why aren't you the liar?
Fourthly, you can always go one level up. Show the footage of them constructing doctored footage.