The naughty username checking system used by Twitch
ghostbin.com
ghostbin.com
I worked on the application form processing for the Nectar card launch in the UK back in the early 2000's, and we had several cases. Luckily we had human data-entry clerks in the loop, so all we had to do was flag when a name contained anything on the "Scunthorpe list" and get a human to look at it. Even then it wasn't perfect and a few slipped through. One of the early PR messes of the launch was someone getting a Nectar card issued with a rude name, and of course they immediately went to the press with it [0]
I'm interested to see they still haven't solved this [1]
[0] I never saw the journalistic interest in this: this guy said his name was <rude name>, and then signed the form to say everything on the form was true. Why is anyone surprised or interested that we accepted his name was what he said it was and gave him a card in that name?
[1] https://metro.co.uk/2015/02/20/woman-refused-sainsburys-nect...
I'm not so sure that has ever been true. Emotional bullying was always a thing, and it did hurt. I'm glad that we're now taking that more seriously.
But it's also true that my emotions are my responsibility, and my emotional reaction to words is unique to me. I cannot force everyone else to be responsible for how I feel about their words. And I cannot expect everyone else to anticipate how I might feel about what they want to say and modify their speech accordingly.
There's a balance in there. I suspect that this balance is what we call "good manners" or "politeness".
Or "empathy". And I think on an interpersonal level this is easy. You can't force anyone else to be responsible for how you feel about their words, that is true. But if you say to a colleague or a friend that "when you call me X, that makes me feel bad, could you call me Y instead?". Depending on the friend they might want to understand why this feels bad, but they would call you Y from then on. Easy peasy.
But on a collective level we aren't yet (if we'll ever be) as developed a civilization to have empathy on that level. So some people want to, with good intentions, extend their empathy to others and write code like this. But seconds later, the creative and determined individual called cVnt_h_ater have bypassed the code and can proudly wear his misogyny on his sleeve.
Why deny him this right? Isn't it convenient to let kids and mentally unhealthy people to identity themselves this obviously? It's always easier when you know what to expect.
I agree. Which is why we have rules around politeness and manners, which roughly conform to same actions as actual empathy. Which makes sense - people don't always have the same level of empathy, there are lots of neurodivergent people who will never understand empathy but who can learn a set of rules that will allow them to avoid social problems.
Or at least, we used to. I feel like this respect for being nice to each other so we can all get along is on the wane.
Or it might be that I moved to Berlin, and Berliners are famous in Germany for being rude/abrupt, and Germans are famous globally for being direct ;)
I think this is depending on how you see it. The modern landscape with social media makes activists very effective in creating information campaigns that raises awareness of how to be polite in a modern world. Eg. nowadays personal pronouns are complicated business but finding out and respecting individuals chosen pronouns are considered polite. A successful change of culture that increases empathy and respect of fellow humans in general. But the downside is that the opinions of the people who don't want to be polite in this way are also amplified and they may want to resist and pushback that triggers them to be contrary and act with less empathy and politeness to make their point. The thing is, positive changes in culture usually last and backwards reactionary people die so that makes me kinda hopeful.
In one way, that is of course true, and people who just blithely ignore that, or even explicitly refuse to use the pronouns they're asked to, are of course arseholes.
But OTOH, it feels like it's getting more and more to keep track of -- at least if you also want to keep up the old norms of "good behaviour" we learned as kids. And perhaps that's precisely why some of us are feeling those old norms are falling by the wayside -- maybe people only have some set amount of attention to spend on this stuff, and if something new is added, something else gets crowded out?
If that's the case, do we know for sure that it's unequivocally a good thing? I'm not quite certain.
In my mind, part of that is that the rules are getting stricter, so people respect those rules less. Over time, the needle has moved further towards the idea that anyone can be offended by anything anyone else says. Even if it's clear (to the average person) that the speaker meant no harm (or even intended a totally different meaning for the words), the fact that someone is offended means that the speaker is wrong. Professors get sanctioned for using a racist word in a discussion about racism (and sometimes that word in particular).
At some point, people start seeing those rules as ridiculous and no longer respect them (though they may still feel compelled to follow them, for fear of repercussions).
What do you mean by this? Are you talking about some contrived example where one person had their parents murdered by a maniac while shouting "bicycle" over and over which causes that word to be triggering ptsd for that individual. In that case that idea is correct since that, hypothetically, can happen for _any_ word yes. But how is that relevant to any of this?
1. A professor fired for using the N-word in a discussion about racial discrimination
2. A professor placed on leave for using a Chinese expression that sounded like the N-word
3. People that express the opinion that someone who has undergone a gender change operation is different from someone that was born that gender. Expressing an opinion that they're different in some ways; not mocking or insulting them.
4. A man saying "I'd fork that" about a software project and everything went sideways because someone else decided it was sexual and was offended by it.
On and on and on. It has gotten to the point that if someone THINKS you crossed an imaginary line, a line that keeps moving, then you can face severe personal and professional repercussions. I do not believe that this is good for us as a society. People are supposed to have differing opinions and then discuss those differences. People are supposed to be able to discuss things in general. People are supposed to be able to have a little fun.
Accidentally offending someone has turned into a horrible offense, and it's bad for our growth as a society.
If you didn't understand that this was what the GP meant you must have led an exceedingly sheltered online life until now. (Which, frankly, feels so unlikely that it feels nearer to hand that your comment was made in bad faith or, at best, unthinkingly.)
But cool story bro. Fight for your right to be a dick I guess.
>>> [@dehugger:] 3. The words you use for this are trans and cis. No one is going to get upset at someone for using them.
>> [Me:] The word the most rabid "PC warriors" use for this are "man" and "woman". If you dare say "Trans women are different from other women in that...", you can bet there's some looney who will jump down your throat [...]
> This hypothetical situation with the looney you made up is not very fleshed out.
Fleshed out enough for anyone reasonably intelligent to get the gist, I would have thought. What, exactly, did you not understand?
> The looney seem to be referring to an individual woman that is somehow out of frame in the story
What "story" -- my hypothetical? Since "the looney" is prominently mentioned in it, how can they be "out of frame"?!?
> and judging by the looneys remarks, the person who now is talking about trans women in general seem to have made previous remarks about the individual
No, absolutely not. Where on Earth did you get that from? My hypothetical was talking about trans women in general, period. There's no need to make up any complicated backstory, because what I said was all I wanted and needed to say: If you dare say, in a general discussion on the Internet, e.g, "Trans women are different from other women in that...", you can bet there's some looney who will jump down your throat BECAUSE you used the term "trans women"; to some people that, too, is anathema because it "singles out" trans women from other women. Some people are so blinded by PC that to them, acknowledging any difference is a sin.
> that was insensitive to that individuals desires to be referred to as a woman and not a trans woman. Which is kind of a dick move...
Yeah, in your weird made-up story that has nothing to do with anything. You are of course perfectly free to make up little fairy tales and publish them on the Internet, but not to impute them as back-stories to anything I (or anyone else) has said. Don't put your shit in my mouth.
> But cool story bro. Fight for your right to be a dick I guess.
If anyone is "being a dick" here, AFAICS it's you.
Purity spirals never end well. The most famous one ended up with people being burnt alive as witches.
But it is a purity spiral - if you don't intend to be considered "pure" then it doesn't matter and you can effectively ignore it. Though whether your employer agrees or not is a different matter.
There are absolutely people who will use this as a reason to put X on a billboard.
Basically, the pen is stronger than the sword, kind of morale.
Sticks and stones may break my bones but words will hurt forever.
After Googling, I think I might have picked up this version from Tim Minchins song "prejudice" [1] although Google also come up with this clip where a senator also use the "break my heart" version, clearly when it was meant for him to say "never harm" [2]
Now it is more "sticks and stones will break your bones because your words have hurt me"
Surrounded by people who used to say That rhyme about sticks and stones
As if broken bones
Hurt more than the names we got called
And we got called them all
Even when the insults could be considered objectively offensive I'm still generally okay with it. I've been called names quite a bit in the workplace over the years and I'm okay with it if makes them happy. Being autistic I don't really care and neurotypical people seem to get enjoyment calling people names so it's win-win.
High EQ individuals have told me this is wrong and the appropriate reaction would be to take offence and try to make them feel bad, but I've argued this would just result in an objective reduction in happiness in the world so it wouldn't make sense. Plus, most of the time my colleagues have been nice to my face so it's not like it's ever got in the way of me doing my job.
By the way in the middle school we used a derogatory for gay every day without even having any idea of what did it mean, all we knew was that was the worst word we knew. When others bullied me (always (or I just never cared and don't remember the verbal) physically, way harder to endure) I occasionally went mad and called them this word repeatedly.
https://www.oed.com/viewdictionaryentry/Entry/67623#eid46436...
As the sibling comment eluded to, you definitely seem to be the "High EQ individual" in this story!
I'm of two minds about this, because of the generality of this blanket statement. Why should we accept this as the norm? I would classify these people ("enjoy calling people names") as abusive, and think that classifying this behaviour as neurotypical effectively legitimizes it.
That said, nicknames are also an expression of affection and inclusivity, and derogatory nicknames are not necessarily meant as an insult. Still, they're just as easily used to demean, belittle or dehumanize the target, so it depends a lot on context. I understand what you mean, but I disagree with your general statement.
Not sure how you win exactly, but others in a similar position but sensitive to name calling certainly don't.
As for name calling and slurs at work, I don't see a reason why these should ever be tolerated. Someone who cannot demonstrate a minimum of politeness towards their colleagues shouldn't be part of the team. To me this is a no-brainer.
Usernames are a bit different IMHO. In a general social network, offensive usernames give important clues that allow you to avoid all contact and block users before you even talk to them. So maybe banning them does more of a disservice than a service to other users.
The problem is when language changes for some but not all, so suddenly words that used to be perfectly normal are considered slurs: Those who learned and use them in the original non-offensive sense are suddenly, and in their own opinion unjustly, seen as bigoted.
You may be wrong, and these "High EQ individuals" right, in the sense that your acquiescence to this bullying behaviour might encourage the bullies to keep it up with others too. In that sense, your standing up against them would be not just for yourself but for all their potential future victims too, which would increase the future sum total of happiness in the world.
(I'm not saying it is necessarily so, but it may be.)
No, verbally threatening physical harm induces reasonable fear and affects your logical reasoning because now you have to consider real risk.
Words are not just symbols on paper, they are image triggers in your brain.
But I think that was Roddenberry's point: we can only advance after excising that primitive baggage.
I mean I don't like the (self) censorship either and don't mind rude or mature words, since we're all mature here and calling a fucking asshole a f*ing a$$h0le is an insult to people's intellect.
"Sticks and stones may break my bones, but words will never hurt me" is something told to children in the playground who are too young to know better.
You might as well argue that there's no such thing as malware, because its entirely harmless until executed on a processor.
Also, I dont see any sense in actually enforcing any of those lists. I always think about the "Journey of Life" in The Grand Tour, which, frankly, is completely inoffensive in any other language than english https://www.youtube.com/watch?v=BLXe2WTYngQ
There's been a running joke on the UK Reddit subs recently about people getting short bans for using the term "faggot" as it's now on some automatic, site-wide blocklist. The bans are completely context free, so people are supposedly getting banned for discussing the food product, a kind of large meat-ball made primarily from minced offal. The same happens with "fag", meaning a cigarette.
I wonder, is it too much to ask that we (as in, the various English speaking places) understand the different meanings of words that are offensive in one dialect but have a different and mundane meaning in another dialect?
Old math prof of mine was named Dr. Cock, and he wasn't the only one, the name was common. Students called him Dr. Octocock.
Fwiw, it's well-known as such, after relatively high-profile incidents involving people with that address: https://en.wikipedia.org/wiki/Scunthorpe_problem
(There's a ridiculously large list of innuendo place names for tiny villages: https://anglotopia.net/ultimate-list-of-funny-british-place-... )
The Register has lots of readers outside the UK, and it has never shied from -- on the contrary, seems it positively delights in -- blowing this stuff up in its pages.
Georgyo has orgy right in the name. More and more places are refusing to accept it when signing up.
You have a customer who gives her name as "Fanny Batter". Is that a permissible name?
What about "Juan Kerr"? Or "Amanda Huggenkiss"?
It's not about replacing "ass" with "butt", but detecting rudeness/humour in context.
I'd be amazed if machine learning can solve this one. It's extremely hard for humans.
* all those names should be rejected btw
No they shouldn't.
And if not, then just tell the press: "They typed in the stupid name, so they got it on their card. How is that our fault? Now go write about some real news in stead of this lazy sensationalism."
It's not that big really, half of it is the plinth to be honest.
Don't know if the GP meant that they had soundex, or that they invented it, but in any case it seems they would have needed it.
Of all the ways I expected a talking banana to backfire, I didn't expect this one. Thanks for sharing
Was the no-links policy there from the start? I think that would have helped a lot. As it is, you're allowed basically one outbound link from your profile, so there are link-expander services which people use to link to more things.
That sums up the WoW Classic (and I'm sure many other gaming communities) a little too perfectly.
The hardest problem with the implementation was that with a long list you can't just search for a few dozen inappropriate words (like the Twitch implementation). It would be very expensive to do hundreds or even thousands of checks against every inappropriate word.
The solution we came to was to truncate all the inappropriate words to either 3 or 4 letters and store them in a big set. We then take our generated strings, which are usually 11 characters, and break them up into all possible substrings of lengths 3 and 4. For example, 1a2b3c4d5e6 would be broken down into 1a2 a2b 2b3 b3c 3c4 c4d 4d5 5e6 1a2b a2b3 2b3c b3c4 3c4d c4d5 4d5e d5e6. An 11 character string would always have 16 such substrings. We then check all 16 against the banned set. 16 lookups into a set is pretty cheap and as we have expanded the word set over time (e.g. add a new language) our performance hasn't changed.
One drawback to our approach is that we do have false positives but we did the math and our space was still large enough, the cost of generating a new one was pretty low, and customers never see it so it's just not a big deal to throw out false positives.
Edit: I'm being downvoted so I want to explain - the internet is a huge mishmash of different cultures and all I wanted to say is that it is allowed to swear because I though that maybe, in their local one, it is not and they think it's universal
If the string is a url, imagine sending https://somesite/wanker to your client, when it actually could also be https://somesite/ay3ugd
It's random, I swear!
And this, ladies and gentlemen, is what it would show BEFORE the filter... but after (runs the code again, and prays it works) ... NO PROFANITY!
"Had we not done this work, that link would have been sent out to one of our users." was very well received.
It's a simple solution. Sure, it is still possible for something to slip through that looks similar to something bad. But the potential to strongly offend is greatly reduced.
Also note that if you're too naive about checking for 'naughty' words, you get https://en.wikipedia.org/wiki/Scunthorpe_problem
What is 'v' in this context?
Edit: thanks for the answers. It makes sense now.
the combo "cv" could then become problematic.
> algorithm tries to avoid generating most common English curse words by never placing the following letters (and their uppercase equivalents) next to each other:
> c, s, f, h, u, i, t
https://hashids.org/#how-does-it-work
E: ah it was already mentioned later on, hadn't got that deep into the comments yet!
Edit: But I'll concede that when your outputs are only four characters long and end users will actively interact with them (write them down, type them again later, etc.), additional safeguards might be appropriate. Or simply omit all alphas and use only numerics.
You're still not out of the park with numerics - people with 1313 or 6660 or 4444 or something will complain a lot. The possibility of a 666 in some new biometric government IDs in my country rose a massive stink from church...
Also, have a feelin you meant to do 1312. What’s the issue with 4444, though?
> When Beijing lost its bid to stage the 2000 Olympic Games, it was speculated that the reason China did not pursue a bid for the following 2004 Games was due to the unpopularity of the number 4 in China. Instead, the city waited another four years, and would eventually host the 2008 Olympic Games, the number eight being a lucky number in Chinese culture.
Thought this was particularly interesting.
> What’s the issue with 4444
4 is pronounced similar to "death" in sino-japanese languages and dialects.
Yeah. An important, long-lived ID that will stick with an individual for their entire life, and that they may want to commit to memory. That seems like a good time to take a hypersensitive approach and adopt some kind of filter.
That works pretty well until you realize that some numerical combinations are common neo-nazi codes and may lead to ... unfortunate associations. The ADL lists a few of those^1, but the list is by far not comprehensive, codes actually differ based on locality, and accidental combinatory collision in a 10-character space than it is in an alphanumerical 36-character space.
[1] https://www.adl.org/education/references/hate-symbols/88
I would have said "why bother" until this happened to us.
A customer rang us up in a fury because some demo/ random data that we generated happened to have the word "penis" in it. They were convinced we must have put it there because we thought he was a cock. It was very difficult to defuse the situation.
[0] https://en.wikipedia.org/wiki/French_Connection_(clothing)
I was doing a student event when I saw someone wear "K1SS MY 4RSE". I told him "What an obnoxious hoodie.". He meekly said "I thought it said 'Kiss my force'.". I later saw him in a corner praying. I should've asked him what Allah would've thought about him talking to him wearing that hoodie.
Aah, the good ole "one customer is unhappy, let's waste a week of time on this" approach to IT management. Takes guts to tell such customers "here's your refund, now piss off", but it is the right thing to do.
Our main concern was whether we needed to increase the size to 26 to account for the loss of keys. After doing the math, a 25 digit random string has a ~5% chance of containing one of 150 three or four character inappropriate substrings. That 5% loss isn't that big of a deal. But we had to figure out the math as part of due diligence before shipping.
It worked surprisingly well when we used it.
It does remind me of the XKEYSCORE (Snowden leaks) that used keywords to bubble up potential threats from emails etc https://www.businessinsider.com/nsa-prism-keywords-for-domes... .
It used to be "if I search for this term, am I accidentally going to wind up getting goatse or something?" The good old days.
Now it's "if I search for this term, is the FBI going to kick my door in?"
you need the domain to make the goat-sex joke work
That's what the linked page says (I'm not familiar with Dutch Wikipedia specifics), but that seems like such a strange statistic to track instead of minimum edit count.
So weird that I looked up how to configure MediaWiki autoconfirm requirements. They probably mean a minimum of one edit. Auto-confirmation considers only age and edit count, and I don't see why clicking 'Edit' is so meaningful a condition to warrant developing an extension.
https://www.mediawiki.org/wiki/Manual:Autoconfirmed_users
Searching about extensions did lead to discovering an obscure feature: there's now a built-in URL shortener: https://w.wiki/Q8
Edit: interestingly, the single-character ones seem to have been pre-planned: https://w.wiki/e, https://w.wiki/E, https://w.wiki/4
In the Dutch WP? Because otherwise I should qualify.
https://meta.wikimedia.org/wiki/Help:Unified_login
You can list your local accounts at
The engineer would write something for every test case the product manager complained about, anything else computationally easy, and call it a day.
I once had to implement an audit logging system. What was supposed to be logged? "Important actions." Nobody on the team could define it. We just logged every database write along with the username responsible and called it a day. Nobody ever followed up or inspected it.
Same deal. Both exist mostly for compliance.
Tracking down "why is the antique system suddenly slow". Power went out, system came back up fine, everything but the one ancient but vital app is fine. Dig, dig dig, there's this old dot matrix printer in another room (because it used to be loud and annoying) that no on has fed or looked at in years.
It finally died with that outage, and it not accepting data was the problem. It had cheerfully printed the ribbon through, then fed out the rest of the box of paper it had, and that might've been several years before i saw it.
The roller the paper was supposed to ride had been eroded. The metal rods the print head rode on had a perceptible bump at the ends of the normal stroke.
The fix was a little dongle for the printer port that held the appropriate "i'm alive" lines up. hardware /dev/null. I'm thinking it was 25 pin rs232 because I remember a lot of cussing over it.
That kind of thing is probably fairly common in the industry.
I'm pretty sure it was the only dot matrix printer with a serial port i ever saw. Even daisy wheels were parallel port by the time this went in; but they had a like 50ft cable to move it to the other room. Someone worked hard and paid large to set that up originally.
It's not a list of words used by the NSA or any spies. https://attrition.org/misc/keywords.html
$Username=<username>
$Password=<password>
Connection.string=($Username, $Password)
The parser would flag this - Password was being stored to a variable! So we just changed our code:
$pw=<Password>. Problem solved!
https://www.youtube.com/watch?v=bJ5ppf0po3k
He made a filter that did analysis of the pronunciation of the messages that were being sent. His breakdown of it starts at around 13:30
I see the same here. It's not clever, but no one has any doubt what words are being checked.
This must be more nuanced. Maybe, it's "additional algo processing when a word is hit", eg another layer before "involve human".
This happens all the time in cartoons that appeal to children but also contain subtle adult jokes so everyone can enjoy them on different levels.
What worked in the end was having any newly created thread send a message containing the post title & body to a slack channel specifically for monitoring the forums. Employees and our forum moderators were in there, and any bad threads were nearly instantly deleted. Eventually the spammers mostly gave up. Hard to beat a dozen human brains :)
Having a dozen reviewers, especially if spread across timezones, would be a dream!
CREATE OR REPLACE FUNCTION is_blasphemy (VARCHAR) RETURNS BOOLEAN STABLE AS $$
SELECT replace($1,'_','') SIMILAR TO '%p(o|0)rc(o|0)di(o|0)%'
OR replace($1,'_','') SIMILAR TO '%p(o|0)rc(o|0)mad(o|0)nna%'
$$ LANGUAGE SQL;c) "Porco dio saranno mica i testimoni di Geova? No eh diocan digli che i signori sono fuori, non ho tempo per stargli dietro." ("Fuck, they can't be Jehovah's witnesses, can they? Tell them we're out, we don't have time for their shit.")
The main 'bad words' filter of English Wikipedia:
https://en.wikipedia.org/wiki/Special:AbuseFilter/384
The page (on en:wiki) listing all filters, which also have uses other than detecting abuse:
https://en.wikipedia.org/wiki/Special:AbuseFilter
The special page has the same title in other editions of Wikipedia and other Wikimedia wikis, though many filters are set to hidden. Dutch Wikipedia, for example:
- should be in its own application with its own rules engine so you dont accidentally whack a bunch of userames
- I would have done in the past and cringe when people ask me to update it.
So yes, Amazon employees have ended up with this particular file on their desk.
I have some fun emails whose subject line is "here are your ISIS family log-in details"
and this was right around the time the terrorist group was frequently in the news
The Iraq War began in 2003.
Bush's administration ended in 2009.
ISIS gained power in 2014, 5 years into Obama's administration.
That's not even offensive and there are still way more weed-nicknames available, so I don't get it.
Now I will never have an IG account linked to my FB. Oh well, I can live without that but thank god I don't depend on that platform for business. It made me laugh but there is zero appeal available.
To our surprise, the end user did not feel offended at all. In fact, they were happy because we responded instantly instead of the usual 24 to 48 hours.
[Story]: https://idiallo.com/blog/do-you-make-your-customers-wait
Why portray the bad word as 3 letters - it doesn't make sense.
Looks like "Baggins" is banned.
Poor Bilbo...
Poor Mike.
The story we know was written by him. I wonder what the trolls would say, or the dead dragon, or the town that his actions helped destroy?
Is he a hero, or, just the guy who wrote it all down?
And just when something really important happens, he throws a powerful weapon(the one right) at his nephew and goes away to retire!
A life lead with riches (gold from the trolls), a ring granting him extremely long life and health, and yup.. off he goes, first sign of real trouble.
Poor Bilbo indeed!
You can still use any other language except English to achieve the goal .
That way they could more easily maintain similar looking and sounding words, including leetspeak and other variants.
Because with this kind of approach, something like "yolocaust" will get through, as most checks only go with exact matches and permutations will always get around it easily...
Whereas with something like a levenshtein distance you could compare it with a set of words and syllables and if it's too similar looking, e.g. 90% the same distance compared to username length, you could simply block it.
Is there a legit service I can use or some actual well tested library that can help me? I’m using node and Go so either languages.
Are you making such a service?
Should optimize checks with a de-obfuscation function (attempt to expand non-AZ back to AZ, even if that then shoots permutations at banned word-runs).
It should probably also look more like a spam scoring system, where really obvious stuff is hard-trashed but borderline things are flagged for review / discussion.
I am also very disturbed that, as with most censorship, 'obviously bad' things such as terrorism/etc are co-mingled with 'is adult' as a negative check.
It seems reasonable for Twitch to have validated 'safe for minors' areas where names are filtered. Generic areas, where things are in the gray area and unchecked. Adult Only areas, where swears, profanity, maybe even some of the hateful things are allowed. Informed consumer choice.
Looks like Twitch doesn't like weed.
In my youthful innocence I had assumed the characters in the movie were trading silly, made-up names (e.g. "Who are you calling "sparky", Popeye?"), not racial epithets.
The fact that this list needs to exist makes me sad, but I am glad technology can assist on some level with the issue.
CREATE OR REPLACE FUNCTION is_hateful (VARCHAR) RETURNS BOOLEAN STABLE AS $$ SELECT ... OR replace($1,'_','') LIKE '%aggin%'
Bob: "Sure no problem if it's only temporary"
--few months later
Chief architect:
"People are spamming more and more, we should design new system for these 234 new bad words, I need team of 7, two backend guys, 5 frontend guys and 4 weeks. It will also require minor rewrite of few external components."
Boss: "geez we're in the middle of sprint right now, Bob can you add these 234 words to existing filter? Make sure it's in production before lunch, thanks" (checks watches) "I have to go now, meeting with customer, bye".
The reason is: This can be updated very easily and out-of-band of other deploys of the main application code.
To make this work it has to be super agile to update. Probably this version we are looking at is an old version that happened to be put in git. The functions in prod probably have several additions since.
Sometimes it’s just not worth the effort to try and help people solve these problems.
- Have people choose any username they want
- Have an “after the fact” human review system on usernames
- If the username is inappropriate, change it to “smallsausage[0-9]+” without the option of reverting it back or requesting a new user on the same e-mail address.
Also, Lisa Pedro is not welcome there.
Looks like "SeeKylePlay" would trigger an insta-ban for "Sieg-Heil". F.
I have a hard time believing this silly mess is an actual component of anything.
sukciD suggiB was here..
Just six seemingly harmless letters arranged in a way to form a word with more power than the pieces of metal which is forged to make swords.
Just a couple of G's, an R and an E, an I and an N....
Would regex be really much faster than checking it against a 1000 or more bad word list?
Also bad word list can easily get updated by moderators as well, I really can’t understand the logic behind using so much regex.
But note: for the most part this isn't using regexs, and to the extent it does, it seems largely intended to make the maintainers' lives easier by avoiding having to represent (and maintain) all the permutations they are trying to match for.
What's sad though is that they're doing many, many passes through the pattern matcher, rather than just building a single big DFA from the whole list of patterns they want to match, which gets traversed in one pass.
I pity those who chose that username based on The Witcher's Nekkers.
Highlights:
CREATE OR REPLACE FUNCTION is_tragedy (VARCHAR) RETURNS BOOLEAN STABLE AS $$
SELECT replace($1,'_','') LIKE '%george%floyd%'
$$ LANGUAGE SQL;
CREATE OR REPLACE FUNCTION is_derogatory (VARCHAR) RETURNS BOOLEAN STABLE AS $$
SELECT replace($1,'_','') LIKE '%retard%'
$$ LANGUAGE SQL;
Now we know we can make as many goergefloyd accounts as we want! *devious grin* *chuckles to self*A hand coded if statement