Mimicking a device is becoming almost impossible
multilogin.com
multilogin.com
I find it's very useful to think of the problem not as governed by technical possibility, but rather by costs.
Good spoofing detection does not prevent fraud, it increases its costs to fraudsters (hoping to render fraud uneconomical).
Bad spoofing detection harms legitimate users and is inefficient against capable actors.
All of this gets much more interesting and complex when we consider privacy (and privacy compliance) to be part of the puzzle - now fraud prevention incurs a cost to the defender as well.
Finally, it's important to consider asymmetric risk/benefit counts. Are you defending against a single major heist, or against billions of tiny events (click fraud)? The trade-off will look different.
How else would you test your SPA against millions of different content-modifying browser plugins, for example.
You don't.
That's why most of them seem to be perpetually banned.
[edit] Now, if someone would build an onion router that ran on people's end devices and served as a proxy network for these front ends, that would be interesting ... but would probably end up getting lots of people blocked, so please don't take it seriously.
These scum companies rely on "engagement", so they can only block people if it's a minority. If _everyone_ participates, they have no choice but to accept it (or maybe change their product so there's no incentive to run these unofficial front-ends in the first place).
On a somewhat related note: I used to write sneaker buying bots and captcha was one of the best things to happen to the bot industry.
Captcha was easily "bypassed" by services that had humans sitting at computers generating tokens. Recaptcha tokens are valid for two minutes by default, so you'd be able to generate tokens up to two minutes in advance and have dozens of tokens available when the product in question became available for purchase. Real buyers had to wait for drop and then spend 15s+ filling out the captcha. Bots would take most of the stock in less than a second. I always felt like this was a fantastic example of a bot prevention mechanism that actually actively harmed "legitimate" users.
Not sure how reCaptcha “v3” (the photos of streets one) works. It could work the same way, but just with more known and unknown (unclassified) pictures? It’s also a lot more harsh if you get it wrong: artificial delays, switching to the “click until no more remain” mode (with more delays), and more. Sometimes, it just refuses to let you through no matter what: https://youtu.be/zGW7TRtcDeQ It’s extremely frustrating.
You could actually call this a symbiotic lock-in. I once read an excellent post on how SEO (despite being a gray area for search engines) entrenches Google's market share because many SEO practitioners specialize in it.
If your ecosystem has become so specialized that most well-adapted users are bots, you have a problem (or not, if you can monetize bots).
Can I ask why you quit the Bot industry? It seems super lucrative right now.
This is the whole point of this article: device detection is generally a false sense of security.
It really depends on what the benefits of bypassing the detection are and if those benefits can be had elsewhere at lower cost. For ticket sales or limited run product sales or similar things, the benefits of buying (and presumbably reselling) the limited item is high, and there's a limit to what you can do to detect humans (and as you've described it can be counter productive). For spam prevention, making it cost more to spam means spammers are encouraged to find somewhere else to spam, which is good for your users (while using your product anyway). But it won't stop everything, some people are going to manually type their spams, and some people will build robots to tap out their spams more like a person; and if you ruin the experience for users (especially new users) that doesn't work either; a network with zero messages is spam free, but doesn't help anyone.
What if you only pop up important information via an onBlur event in a dynamically created div element that is overlayed on the page, only after asynchronously querying a REST endpoint to get information from a database?
How would a crawler even know that this element is shown when you leave a specific input field? How would it know to wait a second while the database is getting the data?
Like it or not, we're well past the days of complete HTML documents being served for a request.
site owners want the advertising that was embedded to be seen and thus make revenue from it.
Slack limits the amount of searchable messages to 10k, so if you save them regularly to local DB, you can search without limits.
All kinds of owners don't sell/show ads, or care about scraping.
There is no need to destroy privacy, waste resources, create terrible ux, gate content…
Things that work for small sites fall apart once you get 10~100 million+ people.
Look, I help do the actual user moderation for a large (10m+) subreddit, and it is a god damn nightmare. Plus, on top of that, we have to do automated bot detection anyways, because otherwise it's just neverending bot spam.
Reddit is actually a good example of having millions of users with no captcha. They “fixed” their problem with user moderation.
You chose to become a moderator, so it must not be that much of a nightmare.
EDIT: Reddit does have a captcha now. My guess is it came along with the other users hostile changes in the new design.
edit: They have for at least the last 6 years based on online questions. Long before the new design.
Automation detection and privacy are orthogonal. The ideal bot detector yields a single bit of information: bot or not. It doesn't need to try to de-anonymize users or anything like that. As for UX, well, spam is a bad UX.
The incentive is money. You can't just wave a magic wand and get rid of the incentive to spam things. Asserting that you can is wishful thinking.
My own experience briefly running sites - bot traffic is 100% trash on almost every metric. Actual conversions, spam, click fraud, other fraud (spam talking about work from home), privacy violations (trying to scrap all user profiles, capturing deleted user content etc).
The real reasons people implement these draconian measures range from inept cargo culting to nefarious business models. If you have a problem with spam, add user moderation and call it a day. That’s just a justification though, there is never a pro-user reason to destroy ux and privacy.
Huh?
I require users to register (little to no privacy) and pay a bit - this results in a both better user experience and is a pro-user reason to destroy privacy. You can also blacklist most of the anon email providers / do SMS verification etc
I’m also curious what your site is. Will you say, or do you wish to remain anonymous?
Ideally, you could get rid of every captcha and just have an open API for everything. Somebody wants to create lots of email addresses to send spam? Remove the incentive to send spam.
Somebody posts thousands of manipulative tweets? Maybe the "feed" model of social media is wrong. Don't assume every random account is equally important. Better just have a simple social network where you can connect with your friends. Or even better, design your society such that it's decisionmaking is more resilient and can't be gamed.
The really hard but worthwile problem is not in tech, it is how we evolve our society along with the tech.
But how?
Unfortunately a new generation has come along that believes advertising-supported business models are their birthright.
I respect that services have running costs and I’m happy to cover my usage, but I have to understand that the average consumer doesn’t give a crap, or is simply ignorant and doesn’t know.
Good luck. Literally nobody has figured out a viable way to do that. At this point, we need to consider spam a fact of nature, like viruses.
> Locking down things takes the fun out of them.
Locking down things so we create a space that actually functions and doesn't succumb to abuse. What you call fun, others would call trolling they'd like to get rid of.
https://www.bleepingcomputer.com/news/security/microsoft-adm...
Switching to a more common user agent recently helped only a little bit. At least that is what I believe, unfortunately I haven't done any measurements. It might well be that it helped not at all.
The irony is also that I'm permanently signed in to more services on this computer than on any other one in my household and that it is the only computer where cookies don't get purged regularly.
Also, I never used the RasPi for any automated tasks or anything else that could be interpreted as bot traffic.
The worst offender is Yandex, which is pretty much unusable because it let's me solve the captcha every few mouse clicks.
Any ideas how make a RasPi useable as a work computer?
I don't have a google account and use a desktop Linux-based OS on my phone. I have to solve the same captchas there as on my desktop.
i see captcha everywhere. and try to avoid when i can. basically if it is a comercial site, bye. but gov sites are using it too now.
from time to time i even get into captcha hellban. which should be criminal since they are used in gov sites. hellban is specially cruel as your entire IP is served the same 10 round of images infinitelly.
On the "fun" side, a bunch of 4chaners were intentionally poisoning the results with swear words.
Maybe the average HN'er doesn't do any serious grocery shopping, say to feed a large family, but there's no way I am taking a cart full of groceries to a self checkout line, I'd rather wait. Sure, for buying 10 items or less, they are OK.
I prefer cashiers who are paid well to do their jobs, rather than subsidizing companies that refuse to pay living wages and just threw in the towel. Try checking out an ALDIs sometime to see how truly blissful being rung up at their checkout lines are.
However, I do occasionally need to get just an item or two. If I'm at a store that has a self-checkout, it seems to be faster for that.
Also, the checkout at ALDI is an impressive demonstration of efficiency. CostCo seems to have a similar philosophy.
Also, i dont want self driving cars. Not until all the social issues are resolved.
Ah… thanks for this second of imagination, that was some quality time.
To really cause problems, you would have to incentivize a large number of other web users to also mislabel the same images.
For SMS verification at least there are quite a few sites dedicated to verify you for a few cents.
AFAIK they buy real prepaid sim cards and allow reuse by different customers on different sites.
The bigger problem is if the site requires you to re-verify your phone for whatever reason (for ex. paypal does when you access it from a strange IP or sth) and the number you used may not be available anymore.
For everyone except the user who now has no privacy, is trivially hacked by SMS interception, can’t create multiple accounts (e.g., to segregate their activities on a chat platform), ... .
This is usually the intent of sms verification. I hate it just as much as you but I admit that there is no better method for preventing ban evasion. When accounts are trivial to create, moderators have no power.
Sounds ideal! Moderators are just power tripping gatekeepers. Community ranking of content is sufficient to get rid of trolls, spam…
Taking away people’s right to privacy just to empower moderators is quite dystopian.
I can buy a prepaid sim card in our Hofer (=Aldi) for 2eur, no verification, no nothing,... they even have the same barcode on, so even Aldi does not know which is mine, and I can pay in cash. I can buy a refill there with cash too, but just to receive an sms, I don't need one, because receiving messages is free.
It's one thing the whole web can take out of the cryptocurrency transaction model. Costs can also be explicitly monetary and not just time-based through captchas.
You also have to remember carriers will track the approximate position of a phone with an anonymously-purchased SIM-card, so better not take it home.
It was a bit of a surprise as I flew through Heathrow and saw simcard packs in vending machines.
I had to set up a web account with the IRS in the recent past and I just skipped the entire mobile validation by having them send the auth code to my mailing address. It arrived in two days and worked perfectly. YMMV.
I am, however, interested in what mobile carrier you use that had a "real" (not VOIP) mobile number and can receive SMS from "short codes", but did not work with the IRS ?
That's unexpected - basically every mobile number in the US is one of the big three or an MVNO operating with one of the big three networks ... can you share just a bit more about your setup ?
I think the IRS may have access to customer names associated with mobile phone accounts to confirm identity, but not all carriers have identity information for their customers, and I'd guess smaller carriers (or privacy focused carriers, if any exist) may not provide that access.
Problem. I don't have any other type of number. So while I added my older GV number to my Steam account, I can't update my number to my newer GV number because Valve won't accept non-LL or Cellular numbers for 2FA. So I have to keep a number I've changed from everywhere else... or not have a number on my account anymore.
And you are performing work without being compensated for it. In many countries this seems to be illegal - you have to pay at least a minimum wage.
* Protocol Testing - where you use servers to generate lots of HTTP/S traffic that is correctly structured to simulate user traffic. Generally, this is focused on capacity and server response times.
* Browser Based - where you need the complex logic present in SPAs and Javascript to accurately create test traffic. This also allows for as-close-to-possible real user response times. This is essentially "headerless browsers" of varying types. This requires more performance test compute to process - so often a combination of both types are used together.
I find that Testing application security is one of the most technically challenging aspects of performance testing. Often some parts of the security infrastructure have to be disabled to allow testing to occur. For example, Rate Limiting by source IP, any form of Captcha, 3rd party (OpenId etc) services have to be disabled - which increases the risk to application availability because sometimes there are components that haven't been tested exactly the way they will work for actual users.
Luckily most the 3rd party services we use are already significantly tested by their vendors - but it is something that I worry about.
The examples of people trying to spoof are:
> People or bots who want to get more elements specific to certain devices, or who want to break out of so-called ‘device ghettoes’ (eg they don’t want to have restricted possibilities due to being a mobile device)
OK, but does the website owner really care if a tiny fraction of people do this? Restrictions by device are usually for performance and ease of use reasons.
(And when content is limited to certain devices for legal reasons, like HDCP, this is accomplished with cryptography, not with device detection.)
> Likewise, some threat actors want to take advantage of the fact that some security measures are not as tight for some devices.
Seems like it would be better to patch the security hole instead? Or else deprecate support for old devices (e.g. stop serving HTTP, only HTTPS). Anti-spoofing seems like a terrible solution to security.
> Who can stop people from utilizing device spoofing if a website cannot show captcha to mobile devices even if some rate limits are exceeded...
Since when do CAPTCHA's not work on mobile devices? And if yours doesn't... switch to one that does?
> or if a company offers specific discounts or products only to some types of devices?
That's kind of a dark pattern anyways.
I mean, the article's interesting, and device detection is (sadly) super-necessary for progressive enhancement, as feature detection doesn't work in every case -- but you can assume honest users in that case. If they spoof their user agent and the site breaks, then the problem's on them.
But it seems a little bizarre to me to put development effort into anti-spoofing measures rather than addressing your actual problem directly. Is there a use case I'm missing where anti-spoofing really is the best or only possible approach?
Also fake reviews.
Who uses multilogin anyway?
I'd bet people try to get around quotas or rate limits by spoofing different devices. Maybe ad fraud as well?
Or even better, real device bot farms https://www.youtube.com/watch?v=X_pRsSM_sXQ
Remember this: https://qz.com/1131515/google-collects-android-users-locatio... ?
A passing birdie told me that it was an internal antibotting sting.
You can't do anything about that. And yes, even WebGL fingerprinting was defeated nearly completely.
Device detection is different from content negotiation. This statement is similar to a statement that ignores the HTTP Accept-Language, and claims that location access is necessary to build internationalized websites.
Try, for example, to disassemble Facebook's APK or disable pinning via FRIDA (https://github.com/frida/frida). It's not exactly easy, and with frequent releases, it's a moving target.
https://github.com/berstend/puppeteer-extra/tree/master/pack...
At some point the only way to not be spotted as a spoof is to run the real thing.
If you think people aren't detecting spoofs like this then you are mistaken. From ad-spoofing detection, to e-commerce bot detection, this is very routine for companies to look at, it's not new it's just becoming more available to everyone.
Spoofers would do better running the real things and learning tricks from poker bots by reading the video output and controlling computer inputs. This is fine by a lot of people as it has increased the cost on the spoofer side.
For example spoofing as part of DDoS is now cost-prohibitive as you either cannot achieve the scale needed or you are too slow to be effective... which makes the market for booter service less viable at their low cost.
For ad-click stuff it wipes out the bottom of the market and forces the fraud on higher value adverts where it is more visible.
For e-commerce bots trying to buy the latest sneakers in sneaker drops this cost is irrelevant as the benefits are huge, but... with enough of the other fraud reduced companies that provide services to protect here can focus more resources here to make it harder.
Similar to how the most effective spam I now see on sites I operate is actually now human generated by cheap human labour (effective meaning "gets past layers of detection that stops it early")... the spam problem for me is effectively solved as it's been reduced so much and humans are slow and inefficient.
This stuff is something of a secret weapon to those who know about it. Because so many developers assume it can't work the companies that master it have a large competitive advantage.
Source: About a decade ago I created Google's main "device detection" platform, as this article calls it (not Picasso, the thing that executes Picasso). It's actually more like an automation detection platform, as it's not a fingerprinting or device tracker, it just tries to separate human operated from automated clients. These days I'm told there's a large-ish team that maintains it full time and has ported the concepts to other platforms like Android.
It started as a 20% project because at that time almost nobody at Google took the idea seriously. Fortunately, my manager was happy to support my experiments. People had the same common (but incorrect) intuition you're displaying here, that any sort of client integrity technique is so easy to work around it's hardly worth the bother. Actually even I believed this to a large extent, just less so than the others. This turned out to be wrong for some not entirely obvious reasons related to the structure of the spam industry:
1. Most spammers are either not programmers at all, or are extremely poor programmers compared to a typical tech firm employee. They can in fact be out-coded.
2. This is because spamming is usually not all that profitable, so programmers who get good can find better and steadier money in the white market. The ones who remain are typically those who live in places without any local software opportunities (e.g. developing countries).
3. Because of this mounting even a not very strong defense is sufficient to corral your adversaries into a shallow economic pyramid, in which a small number of "skilled" people produce tools and services they sell the others, who then run the individual campaigns. This means you are probably not fighting as many people as you think you are. Screwing with the supply chain is an excellent way to wreak havoc on spammers.
When we first deployed the system we spent several months tuning it in what was effectively a running battle with the major Google account sellers. We discovered that the sellers were in turn buying their account creation bots from other people, and some sellers were actually re-sellers. One of the sellers had been using a "raw" bot that didn't embed a browser engine, and thus was knocked out of the market for months as they waited for a new bot to be written from scratch. When that came online there were mistakes in its browser automation that we were able to detect. The developer of the bot couldn't de-obfuscate the JavaScript we used (too hard for them) so treated the platform as a black box, just trying random things in the hope it'd work. We could watch this evolution in real time and block new versions as they were released. After a few rounds of this the seller got sick of it and switched to a new bot supplier. This new bot also took months to complete, and when it arrived it had fixed the bug we were using to spot the first bot, but introduced new bugs the other didn't have, meaning even then it was detectable.
At that point the seller gave up, as presumably paying for the development of all these bots was quite expensive relative to the margins involved. This in turn nuked all the resellers that had been relying on that guy, and blew a hole in the entire Google-oriented spam ecosystem. Spammers had to start phone verifying accounts en-masse, and for most of them it just wasn't worth it (a few switched to using stolen accounts instead of creating them). I haven't been there for years so don't know what the current state of play is, but you do still see public threads crop up from time to time where spammers say they tried to beat the system and couldn't, like this one:
https://github.com/BitTheByte/YouTubeShop/issues/14
If you want some insights into the minds of the typical newbie spammer when faced with this system, try this search and flick through some of the results:
https://www.google.com/search?q=site%3Ablackhatworld.com+bot...
NB: Sometimes people claim they've "cracked" this system but usually they mean they did a bit of reverse engineering out of curiosity. Going further and making a real spam bot that can reliably beat it is a much harder thing, especially if you want that bot to be working with HTTP directly for performance. We never saw anyone attempt to build an HTTP level bot that worked against it in the time I was there. Probably there have been some attempts in the years since.
I would expect most websites to assume that my machine can't exist. But I don't have any problems with captchas.
Not true. What about splitscreen?
I think he refers to these kind of UI features on smartphones: https://www.samsung.com/au/support/mobile-devices/using-spli...
https://www.freehaven.net/anonbib/cache/oakland2013-parrot.p...
@inproceedings{oakland2013-parrot,
title = {The Parrot is Dead: Observing Unobservable Network Communications},
author = {Amir Houmansadr and Chad Brubaker and Vitaly Shmatikov},
booktitle = {Proceedings of the 2013 IEEE Symposium on Security and Privacy},
year = {2013},
month = {May},
www_pdf_url = {http://www.cs.utexas.edu/~amir/papers/parrot.pdf},
www_tags = {selected},
www_section = {Communications Censorship},
}CDNs are blocking even my genuine requests.
Options for Paypal and credit card payments (you can choose
one of the alternatives):
* Send us a copy of an ID, issued by your Government, which clearly
shows your name and picture. The file will be completely deleted
after the verification process
* Pass a video interview with our customer support representative
Options for Bitcoin:
* We don’t ask to verify Bitcoin paymentsBitcoin doesn't need identity to prevent fraudulent payments. It uses math for that.
The legacy banking system has no real, complete solution to fraudulent payments. So instead they bodge on this identity-checking nonsense, which maybe sorta works sometimes, with very high overhead. Privacy is collateral damage here.
KYC for cryptocurrencies is like horse-buggy manufacturers requiring a whip and manure-scooper in every automobile. The horse-buggy industry is very desperately trying to convince you that this requirement is for your own good. And that without it, the terrorists will win.
Sorry but they're not getting my ID. Strikethrough.
I am. The changes to Windows over the last year are designed to do this. For starters, if you install Windows without a Microsoft account (which is only possible if you lie to Cortana during setup and click “I don’t have internet access”), the modals and update flows that pop up after you complete installation represent a dark UX pattern designed to make you create an account anyway. Windows has also been updated several times over the last year to default Edge over your preferred browser (going so far as to actually force the Edge icon onto your task bar AND desktop AND force you to go through a “set default browser as edge” flow). Most recently I was auto signed up for a news ans weather widget (with ads) on my taskbar. MacOS isn’t much better these days because even a brand new Mac is loaded with Apple Bloatware (do you want to pay for iCloud? Apple News? Apple TV+? Fitness? All music? All together in one package? Also here’s a new Finder format which defaults you to save to iCloud and hides your OneDrive).
I'm not German but I also find the "thumb, index and middle" method to be the most natural, curious why people in the Anglo world have a different way for doing this.
I naturally use my thumb to hold down the other fingers leading with index-as-1.
I'm not sure what the thumb-as-1 does for 4 though.
Since the thumb is the first digit on the hand (depending on the way you go) I can see it being a logical choice - but not necessarily the most convenient once you get to 4.
However, if I was tallying, I'd do thumb, index, long, ring, pinky. I have a Chinese colleague, and he always goes pinky, ring, long, thumb.
I guess this is all just culturally acquired, no deep reasons necessarily. Still, as someone else was pointing out, sticking out all fingers except the pinky for the "German" 4 seems hard to do.
Definitely sounds like a cultural thing. In Australia, this is an extremely common gesture. :^)
It's, er, culturally variant, but not in a way that raises eyebrows.
The other day youtube was showing me ads, without exaggerating, every ~20 seconds. This would happen for 3-4 minute stretches before they got less infrequent.
This may not sound like much but when you're watching a 15 minute video they add up.
I looked and looked for ways of blocking youtube ads on my iphone.
Morality aside, what struck me when I looked into this wasn't that I specifically could not find software to block ads (which was disappointing), but rather the larger point that came across as I browsed forums looking for a solution - how hard it is to hack your devices - even android devices.
Most articles ended up with having to root your device, which is fine, but even then the solutions were unreliable.
Curious to see what ends up happening in the long term.
That is probably going to be my best bet though.
Wait, what?
> Who can stop people from utilizing device spoofing if a website cannot show captcha to mobile devices even if some rate limits are exceeded, or if a company offers specific discounts or products only to some types of devices?
Aren't CAPTCHAs shown to mobile web browser users all the time? Is there some law against showing them CAPTCHAs in the Peoples' Republic of WestArctica?
> Who can stop people from utilizing device spoofing if a website cannot show captcha to mobile devices even if some rate limits are exceeded, or if a company offers specific discounts or products only to some types of devices?
The "cannot" is not meant in technical, but rather organizational sense - somebody decided that they don't want to show captcha to mobile users ( one reason might be that the user experience was deemed too bad). The same way they decided to offer specific discounts only to some other class of device.
I thought perhaps the author was going to discuss comparison to real world fingerprints. Here's are a few questions for readers: First, how much can a "device fingerprint" be used to identify a person. Is it identifiying a device, or only the person who is using it. How do we know who is that person. Second, is a "device fingerprint" like a real world one where someone can chop off someone else's finger and, as seen in popular TV/film entertainment, use it to gain entry into some highly restricted area. Third, assuming the answer is yes, what stops the collector of a "device fingerprint" from mis-using it. As long as she can make network connections appear to be coming from a plausibly genuine IP address, how would anyone distinguish a fraudulent user of the "device fingerprint" from an "authentic" one. (For example, the collector of the fingerprint could use it to impersonate the true owner of the device.) At least with real world fingerprints, they are physically attached to our person. Neither copying and re-using them nor stealing someone's finger is trivial. And generally online advertising firms are not in the practice of collecting real world fingerprints; those taking real world prints are often government agencies. We may have certains protections under the law against the government. We cannot make the same claims about "device fingerprints".
As a user, I have not found that very many websites/endpoints that I use require any sort of complex fingerprint. I can retrieve the data I want without using a bloated graphical browser, running Javascript or sending a bunch of gratuitous HTTP headers. I send only two: Host and Connection. I cannot remember the last time this did not work. I never see any ads. As such, I struggle to understand all the fuss about "device fingerprints". This is voluntary data transfer to advertisers. If we send all sorts of data to websites/endpoints every time we make a simple HTTP request, then obviously that data is going to be used for something. The user generally has no legal/contractual control over how the data will be used. For example, the user may see ads as a result. Whereas if we keep requests brief and do not send heaps of gratuitous data, it stands to reason we would see less advertising, and certainly less targeted advertising.
If cookies can be stolen, so can cookies that contain "device fingerprints". "Identity theft" keeps getting easier through no fault of the consumer.