https://github.com/robsheldon/bad-passwords-index
I want people to use it. I gave it the most permissible possible license. Please people, use it. I'll even update it Real Soon Now.
https://github.com/robsheldon/bad-passwords-index
I want people to use it. I gave it the most permissible possible license. Please people, use it. I'll even update it Real Soon Now.
My hat's off to you, good sir.
edit: Oh, "Convert your user's password to lowercase and do a simple string search against the index. No regular expressions are necessary." That makes sense, apologies.
[1] https://creativecommons.org/2011/04/15/using-cc0-for-public-...
https://github.com/timboss/Django-Common-Password-Validation
There are a few approaches that would work better for specific environments. I could make a gigantic Bloom filter for example, with lookup functions in a few languages.
But my goal was to make something that any programmer could use in any environment. Substring searching is a basic programming skill and just about everything supports it. (COBOL74 needs external functions to do it.)
So https://github.com/timboss/Django-Common-Password-Validation is just a a longer (10,000) wordlist, for example. If nothing else, it makes what you worked on more discoverable.
The packing program reads in a file, sorts the passwords into clusters based on their length, maintaining a popularity count as it goes. Then it prunes any passwords that occur less than some number of times in the input file ("4" by default, so it doesn't spend a lot of time churning on unpopular passwords). Next, starting with the longest passwords, it works its way down the list, merging the popularity count of shorter passwords into the most popular longer passwords which contain the entire shorter password. e.g., "password" gets merged into "password1234".
Finally it starts writing the output. Starting with the most popular password, it searches up to N passwords ahead in the list (40 by default), looking for any passwords that begin with the same characters that the current password ends with. So it finds "123456", and writes "password123456" to the output file.
I think the only thing I can do to optimize it further is to re-cluster the passwords by their starting character during the output stage, to improve its chances of always finding two passwords that can be glued together.
It's a crude brute-force approach but it works okay.
I just looked at the code, it's not that bad, I really will release it with the next update.
At least if you wanted to keep it as a long string. I'd still probably approach this as a bloom filter as others suggested since it'd let you have a larger number of entries for similar memory footprint.
The file is small enough, and the approach straightforward enough, that for instance somebody could write a trivial Wordpress or Joomla or Drupal plugin for it, and they wouldn't have to think about Bloom filters or have ready-made lookup functions.
I placed greater value on adoptability than cleverness.