Fake WhatsApp update from “WhatsApp Inc.” with Unicode whitespace: 1M downloads
twitter.com
twitter.com
The proposal was that the TLD administrators should whitelist in non-ASCII characters and generally require that domains are either entirely ASCII or entirely in a subset of Unicode that made sense for their native languages - .ru could allow all-ASCII or all-Cyrillic, .gr could require all-ASCII or all-Greek, .de could allow ASCII plus eszett and the umlauts, and could further require normalized encoding (ü must be FC and not CC 88 75) and consider ö.de and oe.de to be collisions [2], and so on. Weird varieties of spaces, dashes, non-printing characters, accents that are only needed to type Klingon, and so on would never get whitelisted in.
I've always thought that was a great idea, and its a general principal App stores could use too. (Although I realize that app stores don't have as strong a concept of a native language as most TLDs do, which makes it a bit harder)
[1]: Its possible it was https://cr.yp.to/djbdns/idn.html, but I'm not convinced. Maybe it was an earlier revision.
[2]: In German, you spell "ö" as "oe" if you don't have an ö key. German speakers wouldn't necessarily need "ö" and "o" to be collisions.
What's to say, 0 and o can be mixed up and ban that too? How would you handle gray areas? Determining what's fishy and what's not is not a matter of black and white. What if you have a German name and want to set up a site in India with ö.in? This idea creates more problems than it solves.
Policing/banning is never a good idea. Internet freedom is way more important than phishing attemps.
Let's ban everything because anything can be phished if you're smart enough.
That's one way to ensure safety of domain names, at the cost of mnemonics ;)
There is a balance to be made for sure, but in the case of both domain and app names, I'd argue that the harms from fishing outweigh the freedom of expression conferred by being able to register "Αpple.com". While the names of things do carry some expression, things have names primarily so you can tell them apart. Its a well settled moral principal that we should have rules to preserve the utility of names as monikers, and that its entirely possible to construct rules for that purpose that have a negligible effect on freedom of expression. If you disagree, then show me the developed country that decided it didn't need a trademark law.
Your hypothetical German expat in India is welcome to try to convince the Indian authorities that the ability to register können.in outweighs the value in preventing amazön.in from being registered. Maybe there are more Germans in India than I know, or maybe phishing causes less economic harm in India than it does in the US.
App store names are harder, because you don't have TLDs giving you a hint to what language(s) most of your users speak. But you can still disallow mixing of Latin, Green and Cyrillic alphabets; you can still say that if you are going to use the crazy accents used in Vietnamese, that you can't also use umlauts; you can still whitelist unicode characters as needed, so that you don't have a dozen different spaces and dashes for no reason; you can still use other signals to give language hints. And as it turns out, most people running app stores have a large pile of money they can use to curate and maintain good automated rules, and hire people to manually audit names when the automated heuristics think its fishy but not fishy enough to automatically disallow, or in response to complaints.
See, this is another problem. Once you have a phishing blacklist, some entity with big enough pockets will just start influencing the DNS system. It would be a slippery slope going straight to dystopia.
Soon, it will become a giant mess and regulation will impede regular companies/people that would question “Who put these rules together? Our domain doesn’t happen to be fishy”.
Furthermore, how important of a name does it have to be for adding it to these automated rules? How do you quantify that?
Even in an ideal perfect world, I wouldn’t want to sacrifice internet freedom for phishing attempts. DNS is a huge part of the internet whereas phishing costs are neglible.
Every time someone brings an idea about regulating the internet, I just get a repulsive feeling. Stop trying to fuck with the internet. The EFF is not enough.
Isn’t Google Chrome and other browsers already doing this? Phishy websites get tagged and there is a big red screen that shows up - most importantly - allowing users to bypass it if needed. There are times when legit but old websites get banned by these Phishing blockers. I think the solution needs to be more “local” than something on a grand global scale.
Different TLDs whitelist different code pages/codepoints as allowed under their domains. See e.g. https://www.verisign.com/en_IN/channel-resources/domain-regi... or https://eurid.eu/en/register-a-eu-domain/domain-names-with-s... and the linked https://eurid.eu/media/filer_public/8d/18/8d18473b-ed9b-4fba...
Which character sets should be permitted in the .US tld?
The general principal of all ascii or all something else is not bad though. It would prevent certain homograph spoofs.
There are also good reasons to mix language in a name. US state abbreviations might be used to distinguish (e.g.) a diaspora community in Texas from their equivalent in Alaska.
Think of the Spanish word "cañon", for example. It's not four English letters and one Spanish letter, it's just five Latin letters (four of which are in ASCII).
The reason I raise Turkish specifically is that the similarity between the characters presents a potential homograph for a phishing attack in non-Turkish domains. (e.g. mıcrosoft.com). Characters with apparent diacritics are less vulnerable (e.g. öracle.com).
The certificate has expired now, and additionally, browsers now show it as https://www.xn--e1awd7f.com/ . Even HN rewrites it if I type it as a full URL in Unicode, actually.
Edit: As for homoglyph bundles, you just have to do it like EURid and create tables.
But the main protective layer would be the user interface, similar to how HTTPS is handled, either descriptively, with warning colors or both.
https://chrome.google.com/webstore/detail/ublock-plus/kjagjn...
https://chrome.google.com/webstore/detail/ublock-adblock-plu...
https://chrome.google.com/webstore/detail/ublock-adblocker-p...
The last two are exploiting the fact that uBlock Origin doesn't come up when you search "adblock".
There's tons more, just look through the search results for "adblock" https://chrome.google.com/webstore/search/adblock?hl=en-US&_... and results for "ublock" https://chrome.google.com/webstore/search/ublock?hl=en-US&_c...
Note that firefox doesn't have this problem (tons of adblockers, maybe some are fake, but none pretending to be uBlock Origin) https://addons.mozilla.org/en-US/firefox/search/?platform=ma... maybe has something to do with the fact that they show usage numbers on the results page.
The rest of the family has been trained to update their files.
Update it every two months
High-level blocks (domains, TLDs) can be bypassed using, say, dnsmasq, by providing specific pass-throughs.
As for sites that use widely-blocked services: the message is to feed back to them and tell them not to do that. The fact of countermeasures does not mean that there will be no collateral damage. In fact, that's kind of precisely the situation that got us into this mess in the first place: putatively legitimate advertising that isn't.
Yeah, sorry, no. As for reporting sites that break, that's a good thing to do in general, but not too good if you want to use the site now.
My point was that if you're going the blockfile route, you can punch specific holes.
I prefer the term "content blocker" because they're more general than just ads. You can block anything you can write a css-selector for (and because of the arms races[0][8][9], other things like websocket connections too). For example, I hide sticky header bars (thanks web designers), youtube comments, everything that isn't the article on news sites and the menu that comes up when you right click on Medium. There are also pre-compiled lists of annoyances and a list to block social share buttons.
uBlock Origin is the only ad blocker that should exist, arguably every single one of the others is fake. There's plain "uBlock" which is the original project that was effectively abandoned in 2015[10]. There's "Adblock Plus" which is a rent-seeking operation[11] that employs 100 people[12]. There's "Ghostery" which is closed source[13] and up until February 2017 was owned by an advertising company[14]. uBlock Origin is the one you want.
[0] https://twitter.com/gorhill/status/846781439890853893 I'm reminded of anti-bacterial soap
[1] https://github.com/gorhill/uBlock/wiki/Dashboard:-3rd-party-... periodically updating files is for cron jobs, not humans
[2] https://github.com/gorhill/uBlock
[3] https://news.ycombinator.com/user?id=gorhill
[4] https://github.com/gorhill/uBlock/commits?author=gorhill
[5] https://news.ycombinator.com/threads?id=gorhill seriously, thanks man
[6] https://github.com/easylist/easylist
[7] https://github.com/ryanbr/fanboy-adblock fanboy also deserves a lot of thanks
[8] https://issues.adblockplus.org/ticket/1727
[9] https://newsroom.fb.com/news/2016/08/a-new-way-to-control-th...
[10] https://github.com/chrisaljoudi/uBlock/commits/master
[11] https://adblockplus.org/acceptable-ads http://www.businessinsider.com/google-microsoft-amazon-taboo... https://en.wikipedia.org/wiki/Rent-seeking
[13] https://github.com/jonpierce/ghostery
[14] https://en.wikipedia.org/wiki/Ghostery https://en.wikipedia.org/wiki/Evidon,_Inc.
the problem in my original comment has been known since June, and some of these extensions phone home https://twitter.com/gorhill/status/898574880773484545 these are probably for "market research"
You are being grossly unfair to uBO's big brother uMatrix. Otherwise, yes.
Mozilla does a manual code review of newly submitted or updated extensions. So, an actual human being sits down and looks at the code. They'll notice when a fake uBlock Origin is submitted.
With that, they also enforce a rule which Google does not have, that any connection to the internet which is not necessary for the add-on to function (ads, telemetry) have to be opt-in.
This isn't perfect protection, for example the extension Web Of Trust required sending browsing data back home in order to function, which they then sold in anonymized form, which was proven to be deanonymizable last year. But it does take out the incentive to spread fake versions in a lot of cases, as you just can't publish an ad-ridden or trojan uBlock Origin clone.
https://blog.mozilla.org/addons/2017/09/21/review-wait-times...
This sounds pretty cool and reasonable. But extensions still can modify the currently displayed website, right? Doesn't that make it trivial to submit data somewhere? E.g. <img> tag with GET params, as the most basic form of this.
https://twitter.com/gorhill/status/898574880773484545
That's another reason to switch to F-Droid.
They DO seem to have some sort of review process in place:
https://f-droid.org/en/docs/Inclusion_How-To/
See: Application Review Process.
I don't know how exhaustive it is or how effective it is in practice though.
(And for the record I'm an f-droid user)
2) The F-Droid maintainers manually build the apps on F-Droid from the respective code repositories. They will notice when something like that is off. This has to do with it being FOSS.
And if you're wanting to tell me that this doesn't scale, not really, no, but it's the same thing that Linux distros have been doing for a long time and Red Hat, SUSE, Canonical actually do have a crapton of users, especially on the server side.
Every one of those distros gets around volume of desired apps by allowing the inclusion of third-party repos (e.g. current Python or docker) which in turn introduces typo squatting as a vector.
You're asking for app store maintainers to slow everything to a crawl and never get popular. No entity which wants to be successful will do that, corporate or otherwise.
"Security via unpopularity"
If it were popular we'd see just as much malware for linux.
So, this strategy would barely work, as users would only look on the internet for a download, if it's not in this trusted repository and then it's gonna be a really unpopular application. (Theoretically, it's possible for your grandma to go on the internet before checking this trusted repository, but that is really just so much more effort.)
If Linux was 90%+ of the market, getting people to download some stuff with a curl command promising some BS or having people download and run sudo would only need to touch a fraction of users to be highly valuable. That's just a random example off the top of my head. And also because I don't use Linux on desktop so I don't fully know how everything works there.
And if someone thinks regular people would be "too scared" of CLI to pipe curl into sudo sh, remember that people are "too scared" of developer tools in browsers too, and yet Facebook and others have to implement self-XSS protection measures in there, because it turns out there's nothing too complicated in computing when it stands between a person and fulfilling their desire (as promised by a scammer).
I'm sure that's not true for all Linux repositories. But humans certainly do look at all apps going to the App Store.
roughly 3% of the world internet users. Sure, 3% is not much, but it's still multiple hundred of millions of people
The actual number of users is closer to 3 billion, so even if your 3% is correct (it isn't) that's not even 100 million.
That's also assuming that every user of the internet is a laptop or desktop to access the internet, but that isn't case. More and more people are only using a smartphone or tablet, especially in emerging markets.
The netmarketshare stats have been hovering around this for a few months, and all the "global internet usage" stats that I could find were closer to 3.75 billions.
Even then, assuming 3 billion, it's still 90 million users... that's most than the inhabitants of any country in the european union
Also, thinking back to the bad old days and the script-kiddie-eseque of many viruses of the early 2000s (iloveyou, et al), I suspect it may come down to attacking what you know: Windows was more prevalent and better understood so that's what people tried to break.
Not my field though, so all just speculation.
They aren’t cutting edge, but neither are the equivalently priced Androids.
Besides machine intervention, this seems like a company the size of Google could easily help (if not solve) while creating a lot of good will by hiring a few at-home workers to better flag or check in on certain applications. Off the top of my head, this wouldn't require a ton of training and you could even provide burner-like phones for folks to download the apps to and play a bit on.
But if you need to download WhatsApp to communicate with loved ones, you can't get it from F-Droid, right?
Your question demands an answer, for real. As does mine.
There has to be a middle ground. In this case, unfortunately, besides close friends or family who know something about tech, you will have to be the one to compromise and use their apps. Plus, I've seen getting people to move to your communication methods which are less popular usually leads to less talking. Not always, just something I've noticed isn't uncommon. I end up talking to people more if we just use FBM or iMessage or something.
The method the other commenter suggested also works.
But yes, these are anecdotal for WhatsApp, which I presume was just an example on your side.
You'd have to mainly use F-Droid and then make some exceptions for those apps. You could also use Yalp store (which interfaces with the Play Store) or Aptoide to get those apps, if you don't want to keep the Google Play Services around, though I cannot make any claims of these being more secure.
Well, I'm not so sure. For example, if you search, in French, for the correctly spelled phrase "Avez-vous aidé quelqu'un aujourd'hui" they suggest the wrong spelling "Avez-vous aider quelqu'un aujourd'hui", which is a grammatical abomination.
I feel like there's a general unwillingness in some realms to do anything at all if the solution is "manual labor until a better tool is available."
This seems to apply to network traffic, whereas I'd guess what to host on the Play Store would be covered by 512(c) instead. If so, the "red flags" test would seem to require at least some automated content checking.
But all the above refers to copyright, rather than the trademark violation and fraud in the WhatsApp clones.
(Hint: it's not just the I's masquerading as l's.)
It's machine-produced and surprisingly good at revealing accidental/unintentional/evil duplicates, considering how cheap it is.
example.com/mushroom-off-the-hook/
Here is an example of the Django framework's "slugify" https://github.com/django/django/blob/master/django/utils/te...TEXT: Fake WhatsApp update from “WhatsApp Inc.” with Unicode whitespace: 1M downloads
SLUG: fake-whatsapp-update-from-whatsapp-inc-with-unicode-whitespace-1m-downloads
For example: limit app and account renames; when creating/renaming app/account, compute levenshtein distance to all the existing ones and if distance < threshold, make it subject to manual review and make it unlisted before cleared.
Problem is, from my observation, that Google has a culture of hating any manual processes, because they do not scale, so they avoid them, unless compelled by law.
2nd problem is that they have big enough market share that they don't have to care about things that are not convenient to them. Slightly off-topic but in a similar way, Apple can increase iphone price 10% per year and get away with it, because people still buy.
Scaling well means "costs / effort scale like f(x) = ax^p +b " with p < 1. Not scaling well means having p > 1.
https://www.recode.net/2015/3/17/11560334/google-is-adding-m...
So how is this stuff just waltzing past their quality control setup? That one unicode character can't really be messing up the whole system, right?
If this stuff is supposedly moderated, who's actually doing the moderation here?
Unicode has a boat load of security issues. http://unicode.org/reports/tr36/
So, the Kindle app's not on my new phone. Because the validation portion of curation is, ultimately, left up to the individual. And I didn't have time to go chasing around the Web making sure I was hitting the correct/official app store page. I probably was. But I've been well-trained to "pause and check" on such details.
P.S. I now recall, causing further hesitation, the "other apps" sections of the search results and/or Kindle app page, included an Amazon Video app. And that app had the same name listed in its details.
Now, the last I recall, Amazon Video was specifically NOT available in the Google app store. Forcing people on non-Amazon devices who wanted to use it, to have to add the Amazon app store and adjust permissions to allow installing apps from it. At least, temporarily; once you had that or whatever app you wanted from Amazon, you could then adjust your devices settings back to their defaults. Unless/until you wanted to pull an update to such an app -- then, rinse and repeat.
So... I see a weird bit of contact information. And I see it also for an app that prior experience taught me was not available in the Google app store...
And, with repeated stories like the OP, I can't trust the Google app store to be well-curated.
What else can I say? Meh...
All of them have "Suhail Mirza, 500 9th Avenue N, Seattle, WA 98109" listed as author.
But they are the real apps: https://play.google.com/store/apps/developer?id=Amazon+Mobil...
That person’s LinkedIn profile claims "I own Engineering for Amazon's Mobile Shopping iOS, Android and Windows Mobile Teams. I am looking for developers globally. Reach out if you are interested."
The Kindle listing I was looking at shared details with the Video listing.
Despite a fair amount of news browsing, apparently I missed the information that the Video app had made its way into the Play store. Actually, I seem to recall some news of same but also follow-up news that it had been pulled, again, within a few days. (The eternal Google/Amazon competition/strife/"user, you are the product" situation.) This would have been months ago.
So, I'm left uncertain whether I'm looking at the real thing, or an imposter. I'm fairly certain I'm not. But "fairly certain" is not "secure".
At the time, I didn't have a lot of time to delve into this. And I only had my phone in hand, making such an investigation more cumbersome.
I didn't install the Kindle app, then. The moment and immediate need passed, and following up on this dropped down my list of priorities.
Those are also things I look for.
Still seems to be in line with my basic point: On Google Play, it's up to the user to assess the item's legitimacy. At least, so far, Google continues to provide these data points to the user; as long as the Play Store itself isn't compromised.
Keep in mind, some of the items recently in question in the news are reported to have had a million plus installs. Separately, fairly recent news stories have described ways in which third parties have managed to glom onto prominent domains -- particularly those providing extensive user services -- to gain the addressing of that major domain for their own functionality.
All the automated or manual safeguards that Google could enact would never prevent people from pulling a fast one, the old switcheroo, a kansas shuffle on each other because it's just something that we do. And we will use whichever means (technology) available, in whatever way feasible. This particular example looks egregious (or ingenious, depending) for cosmetic reasons, but it's fundamentally an interaction between people however fraudulent. Google is in the business of interactions between people.
We don’t expect perfection, but they’ve at least gotta make it harder than the copy and paste bs that litters the Play Store. ‘It’s hard’ is the worst possible reason to do nothing.
Yes, it may require some monetary investment, but we're talking about $700bn company. They could afford it if they wanted to. If they are not doing it, that means they do not want to.
And maybe they implement this, and calculate hundreds of millions (billions?) of Levehnstein distances every day, but the next day someone publishes the same app but with a germanized name ("Was ist App Update") and fools a couple'o hundred thousand germans. Now the solution is obvious, run the names through Google Translate for ALL languages and calculate the respective levehnstein distances! I'ts foolproof! Shame on you google for not doing it already! Simply irresponsible.
Not true. Nobody fakes random products. It's the top scoring ones that are getting faked - for the obvious reason that this is what people are looking for. If you're not in top N (100, 200, whatever), faking you is useless, you just replacing nobody with nobody (exception may be bank apps, where even faking relatively obscure ones can be lucrative, but let's not get into niches for now). Just scanning against the top ones would kick the floor from under the most current fakers.
And of course you don't need to continuously re-scan the data - you need to scan only once, when the app is submitted or the name is changed. So, in summary, when adding app or release to the store, you need to check its name and description against a list - let's be generous - of 1000 strings and maybe run a basic text classifier if you are feel in very AI mood today. Is that impossible to scale? Nope, it's fairly easy.
> but the next day someone publishes the same app
So your argument is because simple checks are not perfect and do not cover 100% of possible fakery, let's not do anything and allow even the dumbest fakers to run free and fill the store with trash. Does it make sense to you? Because it doesn't make sense to me. Probably you decided since your argument won't be perfect anyway, there's no point to even try for it to make minimal sense?
This is a copout. Nobody ask Google to review the source code of each app uploaded. There are plenty of basic things google could put in place to make sure blatant fraud doesn't happen. But they don't, because they don't care or don't want to allocate resources to anything that doesn't have a high return on investment. And since the competition virtually does not exist...
1. There are fewer of these apps.
2. There's less that they can do–on Android apps can do anything and everything once installed.
Not necessarily. Modern Android devices have granular permissions, just like the iPhone.
https://developer.android.com/training/permissions/requestin...
Your comment is a strawman, anybody can be fooled by these kind of dirty tricks. This isn't about users, this is about what Google is not doing upstream to prevent basic fraud on their platform.
I don't yet know of a good alternative. I desperately want one because Apple kit is expensive. But for now "$500 for an iPad" is the advice that gets me the fewest calls for support.
It would be much harder to fake a github.com/whatsapp account than it is to fake "WhatsApp Inc.". Besides the invisible codepoints, one would easily do "WhatsApp Inc", "WhatsApp Messenger Inc.", "WhatsApp IM" and so on.
Is F-Droid a walled garden?
There's a list of guidelines your app must conform to, and Apple is generally more aggressive about catching violations before app release compared to Google. There are consequences to this, like Safari-WebKit being the only permitted browser engine on iOS. Any other browser must wrap this engine.
The grandparent post is likely pointing out the dissatisfaction that devs express regarding the Apple app store review process. It seems like it comes down to an engineering trade-off. At some point you have to choose between developer experience and end user security.
Same thing with Chrome extensions. Mozilla has solved this same problem in near perfection by just having a few actual human beings look over the code of newly submitted or updated extensions.
Google has magnitudes more money than Mozilla, so they could easily afford to just copy that, too.
There has to be some sort of curation. Algorithms and automation can help with the curation, but there has to be something.
[1] https://blogs.microsoft.com/on-the-issues/2017/05/18/fight-t...
I prefer solutions that offer both, freedom and security. Such as proper application isolation, user review systems (a tough nut, yes) and generally having better reputation/quality signals than just a company name.
Walled gardens just keep small developers out of the marketplace by rising the bar. Now you need to pay money or have a name, so WhatsApp, Viber and similar shit can retain monopolies and keep their users despite being filled with ads and offer less security than competitors. Jabber clients and independent media never get a "verified" badge.
If you want to download "genuine" WhatsApp, go to their website, check their TLS certificate (you can never be sure, they don't even bother to get a EV cert, even for WhatsApp web; https://app.wire.com/ has an EV, for example) and follow the link to Google Play. Software repos are not here to do the job of CAs.
Every time something like this comes about I just get more cynical about the complexity of multilingual systems, or systems with interesting typesetting routines.
It's easy to handle, just disallow < 32 and > 127, which are invalid or non-printable chars (but think about tab/cr/lf). Basically every font you can find can display all of the ASCII printable range without room for confusion.
Instead of making everyone use safe ascii charset for IDs (domains, names like the one in the article, etc.), we go for stupid fuckton of language charsets that cause such problems. All in the name of accessiblity or whatever. And all this does it let people continue living in their language-specific bubble instead of just learning the main international language: english and living happily ever after.
And now people suggest some crutches like restricting the data to some subset of unicode. Never learn.
Regarding china, they chose isolationist politics (just like russia lol), it's everyone else's duty to pull the blanket over to the 'everyone' side from 'china' side.
Google has done a good job with some of their "undo" notifications; these work much better imho.
The proper question here has to be what pisses off more users and to what degree does it piss them off: Having to click "Don't ask me again" once after a fresh installation or repeatedly losing some tabs they had opened in the background, because they forgot about them.
Nobody should "fix" something only because users report it.