The Top 50 Gawker Media Passwords
blogs.wsj.com
blogs.wsj.com
I suspect that others do the same thing, and little weight should be given to the strength of passwords recovered from a site such as this.
And here's one from 2008: http://www.whatsmypass.com/the-top-500-worst-passwords-of-al...
In the gawker leak, the first two characters of the stored hash were not a default salt, they were random salts. As for how they're generate? Well, randomly.
Why do you think that what you said is true?
Your comment that the maximum number of a given salt||hash was seven also threw me for a second. Am I correct that that is purely coincidental? Given the limit of only two characters for the salt (in what set? all printable characters?) and the sheer volume of accounts there is simply some unintended overlap? It just happens that the most it occured was seven times?
Yeah, seven is just coincidental. But, it appears that you are correct in one respect: seven seems a bit high to me. I don't have the time to do the probability distributions out, if someone cares would they do the calculation and check?
EDIT: The salt 'sV' occurred 215 times. sV39Fw5at18zo occurs seven times. Assuming that there were only 300 possible passwords each of which occurred with probability ~.3% (the probability of '123456'), then the probability of seven passwords hashing to the same value is incredibly low. Less than a thousandth of one percent. Does anyone know why this is? Or was it just the case that Scorpion's assumption that the distribution is very non-random is correct?
I think it's time for someone to come up with a radically better authentication mechanism.
Yes, you could have a bot that checked the top 50 passwords against a few thousand or so accounts- but even then, you'd only get one or two matches at most.
What if sites blocked passwords that have been used more than twice already? So, at most, there would be two "123456" passwords- any secure password is more than likely something that no more than one other person on any given site would be using.
[Edited: Whoops, did the math too quickly]
Actually, 3000/1000000 = .003, so a .3% chance.
How about instead just implementing a blacklist of unacceptable passwords?
"I'm sorry, another user on this system is already using the password you have submitted. Please choose a different one."
That's not the sort of information you want to be leaking.
If you are going to have a strength requirement, then run your strength validation routine and deny weak passwords. That a password is duplicated seems like a special case of weakness that is not worth checking for.
i.e. You probably don't want special checks for 'password' or '123456' either, since your strength validation routine should catch these.
http://news.ycombinator.com/item?id=2002805
Many comments there.