As a side note, a lot of these spam emails I get are from Gmail.
Also worth noting if you are trying to evaluate gmail's classification performance that the vast majority of what they think was spam is not in your spam label, it got stopped with a 4xx error code at SMTP time. So you don't really have a way to know the denominator.
I'm relaying through SendGrid as I just don't have that many emails coming from/through my server that it's worth the lowest paid level (there is a free tier) to have to worry about it...
I've been considering setting up a higher end server (compared to the $20/mo vps I'd been using) at a data center and seeing what I can manage as a direct mail host without the relay. But 10x-ing my costs just doesn't feel right for something that will take more time and not generate revenue that I'm not that passionate about.
For those curious, been looking at WildDuck mail which seems like an interesting structure and the features are cool, just not sure I want to go through it all. I've been using Mailu via docker-compose on DigitalOcean for a couple years for all my lesser used domains/addresses, relaying through SendGrid. It works but kind of annoying going through setting up each domain added through the relay.
Ironically, SendGrid is the main source of spam passing through my spam filters; but I can't block it because about 1/4th of emails I get from them are not spam
I reported it as spam and Gmail helpfully asked if I'm sure because I communicate with this person a lot and when confirmed said it will block the sender. Unsure if this will have any impact on the emails I send out myself now.
Rather worrying that Gmail addresses can be spoofed.
> Rule #34: In binary classification for filtering (such as spam detection or determining interesting emails), make small short-term sacrifices in performance for very clean data.
> In a filtering task, examples which are marked as negative are not shown to the user. Suppose you have a filter that blocks 75% of the negative examples at serving. You might be tempted to draw additional training data from the instances shown to users. For example, if a user marks an email as spam that your filter let through, you might want to learn from that.
> But this approach introduces sampling bias. You can gather cleaner data if instead during serving you label 1% of all traffic as "held out", and send all held out examples to the user. Now your filter is blocking at least 74% of the negative examples. These held out examples can become your training data.
> Note that if your filter is blocking 95% of the negative examples or more, this approach becomes less viable. Even so, if you wish to measure serving performance, you can make an even tinier sample (say 0.1% or 0.001%). Ten thousand examples is enough to estimate performance quite accurately.
[1] https://developers.google.com/machine-learning/guides/rules-...
3 days ago "Tell HN: Gmail's spam filters have gone bonkers" https://news.ycombinator.com/item?id=34411009
1 month go "Ask HN: Do you all get spam in Gmail daily?" https://news.ycombinator.com/item?id=34093812
4 month ago "Ask HN: What's happening with Gmail spam filtering?" https://news.ycombinator.com/item?id=32923098
"Ask HN: Is Gmail spam out of control for everyone else too?" https://news.ycombinator.com/item?id=30315116
The uptick I have seen is going from 0-2 spams making it to my inbox to 10-20 spams making it to my inbox. When this has happened in the past, I have assumed it is spammers bypassing blocklists by finding new hosts, or by spammers finding a clever way to beat the filter. Usually after these big upticks, they drop off again suddenly, which makes me believe that it was a blocklist bypass and not a filter bypass (my filter is pretty weak and hasn't been retrained/updated in many years.)
The whole system just sucks as a whole, and feels too entrenched to come up with something better. Even a notify+pull system wouldn't fix these kinds of exploits, even if they would correct end-user breaches.
I'm not sure what they've changed internally, because if they have talked about their engineering strategy for spam detection (which I doubt, since it's probably asymmetric information), no one has shared writings about it.
Nevertheless, I get obvious spam in my inbox now, and important email occasionally goes straight to my spam filter now.
People here on HN have been speculating that they moved to some sort of machine learning model, probably because employees were incentivized to pervert the existing product for promotion purposes by gaming internal metrics to prove they've had an impact.
This includes a marked increase in crypto spam/phishing emails due to the cointracker email list breach - those have pretty much exclusively gone straight to Spam (including those using Google Sheets so it has an official Google sender email).
Again, just an anecdote, and I don't doubt that you and anyone else reporting an increase is experiencing it.
Message ID <9UOejz_TlFksgoyXm9GI5Q@notifications.google.com>
Created at: Fri, Jan 20, 2023 at 9:14 AM (Delivered after 0 seconds)
From: "Girl Shows Girl cast a lookSTART JOIN Muriel (Classroom)" <no-reply@classroom.google.com>
To: XXXXXXXXX
Subject: Class invitation: "Check Join now View gambling Babe amidcustity"
SPF: PASS with IP 209.85.220.69 Learn more
DKIM: 'PASS' with domain google.com Learn more
DMARC: 'PASS' Learn more
Message ID <DM6PR18MB3569050DD20FD0372DA98C9DCEC59@DM6PR18MB3569.namprd18.prod.outlook.com>
Created at: Fri, Jan 20, 2023 at 4:50 AM (Delivered after 3 seconds)
From: hoven patroo <hovenpatrool@hotmail.com>
To: XXXXXXXXXX
Subject: 名梦 t94396350
SPF: PASS with IP 40.92.18.30 Learn more
DKIM: 'PASS' with domain hotmail.com Learn more
DMARC: 'PASS' Learn more
You would think they'd do some basic bayesian filtering. This was stuff we fought in 2002.When I worked in this area of gmail we called this the "russian urologist" problem. How do you correctly classify traffic like this when hypothetically some of your customers want to send and receive messages about viagra in russian? Casual observers will say that is spam but not to the russian urologist.
I hope they are just tweaking things!
My educated guess on why? Lawsuits from political parties, notification of class action litigation against Google and others, union notifications, insurance notifications, and similar emails ending up being caught by spam filters.
The lawsuits are piling up.
sort of interesting, the real competitive enemy of gmail was not icloud or outlook, it was iphone.
And every month or so they vary the numbers and I have to tell the filters to route them appropriately to the junk folder. (And I have to tell one mail at a time because if you try to select multiple with different mail addresses the filter doesn't propose to add it to the filter list).
I've definitely noticed an uptick recently and what is most perplexing is that some seem like they'd be easy to catch - in fact, I set up some Gmail filters to do so and they seem to be working 100%.
Examples like, someone put me on their Farm Equipment account, so I was getting receipts and marketing... Also got on someone's college application, so funding notifications etc.
I have sometimes made an effort to contact the org or person in question... I did manage to change someone's password for a dating site, and changed their profile to "I don't know how email works" etc. When I couldn't reach the person.
"Hi - Grandma here - so nice to see to see you last week and glad your studies are going well. I hope you liked the sweater I sent over.
Lots of love."
Still not sure if they're real and granny assumes @gmail.com must be the right email for her grandson's name or it's some sneaky phish.
I get a lot of personal info too like paystubs and legal docs for other people with my first name. I usually reply with "you have the wrong person - you might want to check" and receive a surprising amount of replies like "how dare you read private email - delete immediately or we will take legal action... " Oh well.
"Congrats CALLU !" in subject goes to spam
Like I alluded to in my post, it feels like these would be easy for Gmail to catch directly.
My best filters target the “opt out / unsubscribe” language people put in their footers. I iterate a few times a week as things sneak in. I’ll never get 100% but the results have been very positive.
Example:
from: runnerup info_GBAQBHFLXV@news.ukgkkwwumjhqu.edu via netorg12672764.onmicrosoft.com
subject: --confrmtion-70346102
content: a link labeled: "runnerup-N0tificati0n"
I can't tell if Gmail can't figure out these are all spam, or if they purposely send me 1-2 dozen every day because I religiously label them all as spam and no one else does. Maybe that's because its super hard to label an email as "spam" using their mobile app! My wife has also noted a large uptick in similar spam emails over the past 6 months.
The subject says "paid" -- "Reminder - You have paid an invoice"
but the email says to pay it.
"Please pay your invoice
Coinbase would like to remind you to pay invoice xyz.
Amount due: $599.00 USD"
With sender email being paypal.
1) To get the best training data, you sometimes need to let things you've classified as spam into the inbox to verify that the user marks it as spam. It's pretty standard for training a classification system to occasionally pass negative samples to verify their negativity.
2) The spam filter itself almost certainly has a latency budget, and if it can't respond in time, the message is passed unfiltered. In other words I think the spam filter fails open. It's probably just been down more lately.
Lately it's been Google classroom invitations from sex bots. Along with the random crap that doesn't make any sense, and the McAfee/Yeti Cooler junk.
I have anecdotally seen slightly more types of scam/phishing messages slip through the filter in recent weeks, but I assume it'll go away in the next round of updates from Google's side.
From: ConGraTuLaTioNs!! <info_rGoJOQymodm@appeme.website>
Subj: *robert; $2,613,527 WiNNer ANNouNcement!*
not already detectable by the spam filter? I get them pretty much every dayI always mark these as SPAM, and the next day more recruiter spam comes in. There's no way to unsubscribe because they always are from some random independent company.
To maximize the ridiculousness, Google sends me an email thanking me for each image abuse report or chat abuse report done in Photos -- but they don't seem to be actually /doing/ anything about it.
If any good can come of this long term, it would be the ability for me to charge people to get an email into my inbox. This has been proposed multiple times over the decades, but has never been more needed or feasible than now.
FWIW I have my own domain and switched to Google as backend long ago, and yet I still occasionally have this problem.
A real person who I know in real life, whose messages I care about posts to Google Group from a Gmail account, and the message ends up in my own Gmail spam filter.
Like - the message didn't even leave the Google infrastructure and it got tagged as spam?!
TL;DR: almost all the people who care about quality at Google are gone or not in a position to improve the product
that's as simple as that
No one at google wants spam.
It's going to get even harder as spammers use ChatGPT like tech to write individual spam messages for each person
My point from the higher level comment is that the customer does not care how hard it is. If chatgpt makes it harder there is nothing stopping Google from innovating and improving their detections. The comments are calling out that they seem to be falling behind the curve as more dangerous phishing and spam/fraud emails slip through.
I for one have no sympathy. Google did the same as other giants and gobbled up as much tech talent as they could only to layoff thousands later. If you are telling me I need to feel empathetic for the company reaping trillions from invasive data harvesting and monopolizing the most used digital services on the planet, I shall play the smallest violin I can find.
I look at it a different way, spam and filters are locked in an evolutionary arms race and at the moment spammers have found an adaptation gives them an advantage. In due time the anti-spam filters will adapt as well. It has always been a difficult problem.