I agree. They're voluntarily adding fingerprints to images so I expect the default voice is intentional and it wouldn't surprise me at all (though I have no evidence of this) if the output text has a fingerprint stenographically embedded in it.
I think that there could be even more then a fingerprint in those messages.
It’s this video: https://youtu.be/5CKuiuc5cJM?is=9VQ1FCxY_X3eNm-b
Warning: They anthropomorphize a lot in this video, but I get it… the words exist, why not use them.
Basically we optimize the models to produce output with certain characteristics but that doesn't mean that what we see is the whole truth or even that the relationships in the underlying system are structured in the way that we might expect.