> The only thing that matters here is 2nd preimage resistance, which still has 128 bits of security.
Nope.
Let's assume I have root on the Threema servers and have fully pwned their infrastructure, and want to substitute someone's public key to attack the app. (Remember: The minimum threat model for end-to-end encryption is that the server is evil.)
Alice pushes (3ma_id, pk) to the server, and chats with Bob legitimately. Bob suggests Alice talk to his drug dealer, Dave.
When Alice goes to talk to Dave, instead of (3ma_id, pk) the compromised server sends Dave (3ma_id, pk') where (SHA256-128(pk) == SHA256-128(pk')) and the sk that corresponds to pk' is known to the attacker.
In this situation, a visual inspection of the fingerprint will reveal nothing amiss. The QR code will still validate.
This attack can be leveraged in both directions to make Man-in-the-Middle possible, and the fingerprint mechanism will fail.
The full set of keys and substitutions here looks like this:
Alice -> Bob: (A_id, A_pk)
Alice -> Dave: (A_id, M_pk1)
Bob -> Alice: (B_id, B_pk)
Bob -> Dave: (B_id, D_pk)
Dave -> Bob: (D_id, D_pk)
Dave -> Alice: (D_id, M_pk2)
SHA256-128(A_pk) == SHA256-128(M_pk1), A_pk != M_pk1
SHA256-128(D_pk) == SHA256-128(M_pk2), D_pk != M_pk2
Now Alice's communications with Dave are tapped by a MitM (the attacker) and the fingerprint fails to stop it. And you only need a collision of a partial SHA-256 hash to pull the attack off. You don't need a preimage attack.