Nothing but respect for the engineers at Youtube working on fighting back against it. Seems like a very difficult classification problem.
Nothing but respect for the engineers at Youtube working on fighting back against it. Seems like a very difficult classification problem.
There's probably a dozen or so templates that many of these spammers and scammers use, and if you can spot them in a second, it's not hard to train a model to recognize them also.
It is out of control because YouTube does not care to fix it, not because it's some insurmountably hard problem that one of the top AI companies in the world just can't figure out.
something Minsky vision something (I know it's an urban legend but still, it figures.)
* Cost to build them.
* Cost to keep them up to date in a highly adversarial environment. Note this might mean scraping the solution and starting again.
* Cost of running them at huge scale.
Particularly important is to understand the attack surface is so high and cost to the opposition is so low that people will defeat your approach recreationally just to spam obscenities.
I'll explain too; Money. YouTube make money from the spam accounts, As twitter does, Reddit too. Spam is a very profitable market.
Why else has Google not cracked down on spam accounts? It has the tech, it knows it has a problem and the best remedy they can come up with is shortened YouTube usernames? There's more to it behind the scenes.
A real example is that I worked for a business where my job was to sell email addresses gathered from soft-core pornography websites. A batch of working emails could easily fetch £500.
Reddit will sell you accounts for your "campaign" if it profits reddit and it's all the same for any big walled garden.
You wouldn't even need ML, for the majority of them.
There is actually OS tool[1] that can do this for individual creators. That tool, finds a lot of spam, that YT itself does not, and it's developed by a single guy.
So there definitively is plenty of opportunity to pluck the low hanging fruit.
But Goolge has probably 100 of PHD's developing some kind of ML, that ends up performing worse than some simple regexes, because you don't get promoted for simple solutions.
I comment pretty regularly on youtube to get random thoughts out of my head about the video I'm watching (I'm under no delusion that someone relevant will see my comment), and all the spam replies I get use unicode numbers in the username to share phone/telegram numbers. Circled numbers [0] are a favorite, though not the only ones they use. The Scunthorpe problem doesn't seem very relevant here.
I don’t think it’s entirely true as hopefully the finance department has folks smart enough to realize if headcount is adding negative value or positive value.
The space of all unicode characters is too big for you to look for every single possible "looks like" pair.
Unicode, as unpopular an opinion as this is, should only be used for presentation and not storage. Anything past ascii is a security vulnerability.
While at it, we should also ban the letters I and L as they can be confused. People who have those letters in their name can just request a name change. /s
Leave Unicode alone for the vast majority of the world whose language isn't actually writable using ASCII.
Notice some keys missing there?
The difference is that we have a few good quality fonts, like liberation mono, which let you tell the difference between l and 1 and 0 and O. There can't be a font which does that for the whole of unicode space because a and а are by definition the same glyph.
How would one write an ideographic language using ASCII?
Context, but...
> How would one write an ideographic language using ASCII?
... We're talking about security sensitive context. Users can write stories with all the Unicode they want, just don't use it for identifiers like user and channel names. Reddit got it right.
Because of course this band from "part of the world" has the same name as "other band from a different part of the world" or even from the same country and they have nothing to do with each other (because why google before naming your band - but I digress)
Every spamfilter I've ever used that's worth a damn does that. Sometimes you get slightly odd syntax (I once configured one that replaces A with a 4 so it could detect common fake uses), but this is a solvable issue.
and use the secure subset of UTS 31.
you can eg use my libu8ident, and please dont use the broken confusables list.
The majority of the above sentence is not valid Latin characters for example.
>Mole is too close to Moie, please select another user name. Your bank account has already been charged.
for i in users:
for j in users:
if same_as(i,j):
reject(i)
else:
accept(i)
That is literally the definition of n^2.YouTube and Google are notoriously unhelpful when it comes to instances when they catch a human in their automation drag nets. They lack the capacity to support their platforms so they are probably being cautious not to create support and ill will they can't handle.
Having a person keep an eye on the top 50 popular channels and spotting the patterns in commentor names would be very effective, but having a human is not Google
Incidentally LTT just commented on the announced changes and rants about the prior situation at 15:20, for some insight from a Creator's perspective.
I'm betting it's probably not hard the spammers and scammers to switch templates if old ones get caught.
Yet I haven't seen ANY progress on either platform. Assuming that they are much more competent than a random dude like me as a whole tech company, the only possible explanation is that they have incentive not to prevent spam.
Perhaps: spam = more notifications delivered = more app opens = more potential feed engagement since they've opened the app even if from a spam notification = more ads shown = more revenue = happier shareholders short time... all at the expense of losing long-term trust to the platforms.
False positives should ideally be zero yet algorithms don't have to ban right away.
If there's a relatively small team of people reviewing those spams manuallu after the algorithms flag, I don't see any problems with the approach.
I of course get there are many different corner cases, Unicode characters that look like other letters etc. to bypass the filters yet they're also very easy to implement a filter that would probably detect many spams right away even with many of the corner cases. But again, a human team would still review these tweets anyway giving no room for false positives.
The question is often "does anyone know how to pirate videos?", "...how to invest in crypto", "...how to get fake instagram followers", life coaches, etc.
It's not really custom messages with special chars, it's something that I could detect by using Ctrl+F in the browser. It's not something they'd need special tools on their end to catch. The only thing that is changing is the name of the users.
So yes, I'm gonna agree with GP that "the formats of the obvious spam" are super easy to detect.
I don't know, I've been seeing tons of low-effort comment spam on YT for months if not years. The kind where the comment says "read my username" and the username is "text me on telegram at xxxxx" where x is one of the Unicode numbers with a circle around it.
If only YT or their parent company had access to some kind of spam filtering expertise, built up over decades of real-world use...
What do they mean by special characters? The example they give is "¥ouⓉube(emoticon)". That doesn't include characters with diacritics.
But if this means they ban channels that are actual human names like „Gică Raț“, then this is a terrible move.
People in the comments are suggesting the same should be done to comments. So, for example, people won't be allowed to talk about prices in Japan ("this costs 100¥").