Besides the humor of how Lyft's filter flagged traditionally Caucasian names like "Cummings" –
in addition to the usual issues with non-Western names, e.g. Pimpong and Poon – I'm fascinated/confused how this made it into production? The user database
already exists and was currently being used by the live application. Before deploying this new filter onto the production database, wouldn't you do a dry run to get not
only the count of users who will be flagged and notified, but a listing of frequently flagged names? Which you could easily manually eyeball to make sure there weren't obvious false positives?
With a userbase as big as Lyft's, I'm sure there were a ton of obvious true positives (anyone named "Fuck", ostensibly). I just can't believe they didn't notice a surname as relatively common as Cumming/Cummings. Yes, the apparent naivety of the regex is an issue, but this seems like a system for which a lot of testing on actual data would be easy and natural to do as part of the QA process.