It's 1 in a million chance of any collision across the entire dataset, not 1 in a million records that will have a collision. This is why the article mentions the birthday paradox -- it's taking the chance of a collision between two particular values (and knowing the bits of randomness is enough to give you this) and calculating the chance that there's any collision at all.