How to get gmail.com banned (2011)
mailinator.blogspot.com
mailinator.blogspot.com
I hadn't read that in many years, and what fun to do a re-read.
Thanks Internet - don't stop being you.
I hope you don't mind that I wrote a quick one-liner to see if you're still detecting bots...
@bobmail.info
@zippymail.info
@thisisnotmyrealemail.com
@spamhereplease.com
@safetymail.info
@suremail.info
@mailinator2.com
@spamherelots.com
@mailinator2.com
@spamhereplease.com
@spamherelots.com
@spamherelots.com
@mailinator.net
@mailinator.net
@mailinator2.com
@mailinator.net
@mailinator.net
@mailinator2.com
@mailinator.net
...
Yup :)I didn't see any "evil" insertions, though...
We were really annoyed that rather than just ask us, they had launched what amounted to a DDOS attack. So we thought about how we might exact vengeance...
After a few hours we figured out a pattern to the rogue requests that allowed us to filter them, despite their efforts at stealth (like, they cycle through a list of various user agent strings to make it look like there are multiple different users). We toyed with the idea of, rather than outright banning them, making our pages sensitive to their presence, so that when we detected them, we'd display a false price, defeating their whole operation.
We finally just decided to take the high road, temporarily banning any rogue IP addresses we detected (we couldn't make it permanent because many of the requests came from the Amazon cloud, from which we also receive some legitimate requests)
EDIT: you wouldn't think that requests for a few hundred thousand products would amount to a DDOS, but the bot was rather poorly written and grossly inefficient in the way it walked through the list.
Btw, did you actually return incorrect price data, or did you just insert random bytes, etc.?
How many of you would have an outright revolt on your hands from your QA/QE folks if you banned mailinator? I think everyplace I worked would experience this same issue if we did this.
I used to have a first.m.last@university.edu address and that one was touch-and-go as well due to the fact that the mailbox had two .'s in it. I actually had to file a support request to get Amazon Student to accept it, even. Nobody from a university with that scheme ever registered before?
For the record, the gold standard for email validation is "send a confirmation link and see if they click it". Don't try and get fancy.
One other trick is that Gmail ignores .'s in addresses entirely. first.last@gmail.com is the same as firstlast or f.irstlast.
That means validation was working correctly. The email addresses are invalid even if the server will accept email for them.
Yeah, I'm with the other guy, regardless of whether or not it's a good idea to do validation (it's not), that's not an address that should pass validation because it's not a valid domain or hostname.
I could see it being less of a big deal in the mailbox portion given that it's now kinda kosher to ignore dots there.
[1] docomo customer with two dots in email
Ask me how I know this.
But make sure you have some sort of rate limiting set up, so malicious users can't take advantage to spam someone's mailbox (and get your server blacklisted).
I'm not aware of any RFC that says that mail sent to a+foo@example.com should go to the same mailbox as mail sent to a+bar@example.com (nor am I aware of any RFC that forbids this). I thought that GMail made up that feature and other vendors followed suit since users find it handy.
> Subaddressing is the practice of augmenting the local-part of an
> [RFC2822] address with some 'detail' information in order to give
> some extra meaning to that address. One common way of encoding
> 'detail' information into the local-part is to add a 'separator
> character sequence', such as "+", to form a boundary between the
^^^^^^^^^^^
> 'user' (original local-part) and 'detail' sub-parts of the address,
> much like the "@" character forms the boundary between the local-part
> and domain.
(Highlighting by me)The RFC even gives an example using the hash:
> o A message addressed to "5551212#123@example.com" is delivered to
the voice mailbox number "123" at phone number "5551212".There is no RFC that requires this behavior. Subaddressing within the local part is recognized as a common practice (e.g., in RFC 5233), but nothing requires a system to support subaddressing, or requires a system that does to support a particular separator character or character sequence (e.g., "+") for subaddressing. Email systems are free to implement or not implement subaddressing, and to use any character sequence they want as the separator.
http://mailinator.blogspot.com/2014/10/mailinator-launches-p...
It took me a bit to get my head around the use cases. It's sometimes amazing how many different ways you can twist a simple (complex really) thing like email into a product/idea.
However tricking site scrappers may not work perfectly if the site scrappers maintained a list of websites in their "whitelist". Say if I am scrapping mailinator.com for domain names, if I see gmail.com or yahoo.com, I might just not put them in my database because they are in my whitelist.
Unfortunately it does not work very well as I was not scraping mailinator, but still somehow got IP banned. Fortunately my ip has changed. But they definitely have some strange and overzealous method now.
I would go one step further and look for {spam_words} in "username+{text}@{googledomain}.com", where spam_words can be "junk", "spam", etc. This is like a very narrow edge case, but still might catch something. Again, if you're into that kind of thing; I'm quite skeptical that it brings any value.
http://gmailblog.blogspot.de/2008/03/2-hidden-ways-to-get-mo...
PS: The downvoters are downright silly on this website lately.
OCR requires a lot more programming effort compared to a text-based content scraper
I think that's the basic idea. He could spend his time making it harder to scrape, like the bar across the steering wheel. Some people would be deterred, others wouldn't, and time would be wasted all around.
If the website is doing things right, they have other means (like a CAPTCHA at the least, or phone verification, or you buying an item from them) before deciding that an email address really is an identity.
I realize this was a non-comprehensive list and I'm not trying to just attack it. I think I agree with the core assessment around what constitutes an identity. But short of some really draconian methods, I think you're basically trading off one insufficient method for another. And at that point, you may as well focus on making things easy for people, which typically means just working with email verification.
FWIW, when faced with Mailinator abuse I resorted to requiring a credit card number to sign up for a trial of my SaaS product. The abuse stopped immediately. But there were other impacts to the business as a result. I still debate the wisdom of it and how much of this should have been foresight. As a bootstrapped company, dealing with abuse was just a resource drain and forced me to focus my efforts on dealing with a segment of the population that was never going to give me money. Suffice to say, it was all very disheartening.
Anyway, thanks for sharing your thoughts on the matter.