For reference, a bloom filter is an extremely space efficient, probabilistic data structure that acts a bit like a set and can answer the query 'does the bloom filter contain this entry'. The bloom filter will respond with either 'definitely not' or 'possibly/probably' depending on how it is tuned.
You could conceivably automatically populate this (still hardcoded) bloom filter by doing a brute force language corpus search for heavily correlated word pairs that have one or more of the two words having phonetically similar misspellings. E.g. 'sea' and 'breeze' would be heavily correlated. 'Sea' has a phonetically identical misspelling 'see'. You could then automatically add 'see + breeze' as a spurious pairing to the filter.