I'm wondering why the article fails to mention that there is a sufficiently good and easy mechanism to compare and verify the new safety numbers. You just talk to your peer and read the numbers - and the peer can verify them.
This will fail when AI software gets really good at imitating voice in real-time during casual talk, but we're not there yet (or - if that is my threat model, I'll find an out of band way to verify)