I'm also curious about this. Randomness doesn't actually look very random. There's a surprising number of repeated characters.
It's maybe not so counter-intuitive as the Birthday paradox, but maybe already showing how bad our intuition is at grasping randomness. Initially I had hoped about research on that, which is probably very hard. - Or not? Couldn't you look at the n-gram distributions of the data and look how much that is "random", i.e. are people avoiding clusters harder than they should?