Modern Anti-Spam and E2E Crypto (2014)
moderncrypto.org
moderncrypto.org
(not because I want to dismiss conversation here - did want to read up on it myself)
Also, by the by, there is some interesting context about the author (Mike Hearn) and his recent activities that are not related to the topic of spam.
Warning: off-topic digression below.
Mike Hearn became one of the most visible Bitcoin core developers, working in that community for 5 years, until a well-publicized departure where he declared Bitcoin a failure[0].
He then joined R3CEV, a startup venture that is building a private blockchain platform for a consortium of 70 of the world's largest banks. Hearn's departure was criticized by members of the cryptocurrency community, such as Bram Cohen, who famously called his exit a "whiny ragequit" [1].
Regardless about how one feels about the internal politics of the bitcoin dev community, the R3 project is technically interesting (to me, at least) because it uses Kotlin [2], a JVM-based functional language, and also because it has some interesting design approaches that depart from the established bitcoin blockchain model [3].
I still think that private blockchain platforms have an uphill battle if they want to compete for developer attention with rapidly evolving broad-based platforms like Ethereum and Bitcoin, but the R3 Corda platform is nevertheless worth tracking.
[0] https://medium.com/@octskyward/the-resolution-of-the-bitcoin...
[1] https://medium.com/@bramcohen/whiny-ragequitting-cab164b1e88
[2] https://twitter.com/hhariri/status/790077263572299780
[3] https://gendal.me/2016/10/25/r3-corda-what-makes-it-differen...
> vintage: 2014.
Email is from the 70s and you label anything from 2014 related to it as vintage?
http://www.bitcoinwednesday.com/nsa-gchq-tapped-security-sys...
The debate about permissioned vs. permissionless blockchains ties back into Hearn's thoughts on privacy.
It will be fascinating to see how his work on Corda develops compared to its permissionless competitors.
https://en.wikipedia.org/wiki/Functional_encryption
This would allow a client to combine a server-provided function that calculates a spam score with their private key such that the resulting function calculates a spam score on encrypted email. The client could then hand that function back to the server so it can perform server-side spam detection.
There are a number of drawbacks, including performance and general questions about the security of such a system. That said, I think this is probably the biggest problem (from the OP):
"The third problem is that spam filters rely quite heavily on security through obscurity, because it works well. Though some features are well known (sending IP, links) there are many others, and those are secret. If calculation was pushed to the client then spammers could see exactly what they had to randomise and the cross-propagation of reputations wouldn't work as well."
Using functional encryption to provide server-side spam detection would still require handing a spam scoring function to the client so they can apply that function to their private key and hand the server a result. This would expose the internals of the spam detection routine to all clients, including spammers.
A difficult tradeoff.
Did you mean fully homomorphic encryption? (https://en.wikipedia.org/wiki/Homomorphic_encryption#Fully_h...) The server can compute the spam score under the encryption of an email, and client side decrypts and sorts it from there, so not even the server knows if a given email is spam or not. Of course, not that FHE is feasible, but perhaps this special case is...
The functional encryption scheme only requires a client to bootstrap it. Once the client has calculated the appropriate function based on their private key, they can give it to the server, who can thereafter apply it to incoming emails regardless of whether the client is online or offline.
It's interesting that encrypted messaging has exploded since 2014 but spam has not yet become that much of a problem.
Does anyone have a sense of how difficult it would be to create a service that scans your gmail spam folder and categorizes the contents into 'definitely spam' and 'maybe spam'? I'm probably somewhat of an edge case, but I get over 100 spam emails per day in my spam folder. Almost none make it through to my inbox. However, every month, one or two legitimate emails land in spam. Usually these are there for an obvious reason - "cold call" emails from advertisers and that kind of thing, but still stuff that I want to see, and obviously (to a human) not on the same level as penis enlargement garbage or whatever. (Although occasionally there are real head-scratchers, like I thread where I've already replied to someone twice, and then their third message goes to spam. I guess gmail's filters only operate on the current message and don't look at history.)
Anyway, I get enough false positives that I need to scan through the thousands of spam messages I get each month to try to find them, which is obviously a huge waste of time. If something could go through there, identify the least spammy fraction, and label them, it would save me a ton of time. It would be lovely if gmail offered this themselves, since they already have the spam score for each message. But barring that it seems conceivable that you could do it with a browser add-on, perhaps using a neural network. Still seems kind of like reinventing the wheel though, since you'd basically be building a reverse spam filter. So I'm wondering if there's an easier way...
Machine learning requires large data sets and training, and mushy targets like your inbox are tricky because it's hard to tell a computer what its score was on a given attempt- even with human scoring.
I now run rspamd on my own server, which does a pretty great job. With properly training the bayes filters it has, I now receive on the order of 3 spam messages per day in GMail. Actually, rspamd seems to have fewer false positives than the GMail spam filter -- I guess this could be because it has more information as the original receiver, though?
Getting these results did take some very limited tweaking of the rspamd configuration; I lowered the treshold for what's "definitely" spam (that is, just gets discarded), and I bumped the weight of the BAYES_SPAM rule.
I ask because I had the same problem, with the same setup (own domain forwarding to gmail), adding SRS (besides the obvious SPF/DKIM/DMARC) has really improved things for me.
I'm not sure why it would improve things without also setting up a spam filter? In that case, you're just lowering the reputation of your own server by passing on a lot of spam while acting as if you sent it.
I don't know the rules by which emails are judged, I'm just saying what I've noticed.
Before SRS, I occasionally got legit mail marked as spam.
After that, I don't think I've ever had one marked as spam.
As I said previously: with my own domain, forwarded to my gmail.com address, no spam filtering at all on my side.
Maybe they can tell it's SRS and skipping some penalties?
EDIT: I'm also monitoring my domain at https://postmaster.google.com, but I'm probably not reaching significant traffic thresholds because I've never seen anything other than:
"No data to display at this time. Please come back later. Postmaster Tools requires that your domain satisfies certain conditions before data is visible for this chart. "
Interesting your mention about being grey-listed too. How did you determine that happened? Presumably the same thing could happen to me as well.
Guess I should also check out SRS as mentioned by emilburzo.
Of course, rspamd lets you tweak the thresholds for these levels. For example, after a while I lowered the threshold for "spam", increasing the amount of stuff that gets discarded by rspamd, because I noticed that rspamd was doing a pretty good job of scoring, and the false positives I was seeing had a lower score anyway.
I'm actually not sure grey-listing is the correct term, but I noticed in my server's MTA log that Google was rate limiting me a lot because my server was sending through significant amounts of email. This was also noticeable sometimes because it would take quite some time for email to get through, which I found annoying.
Yes, you probably want to run SRS as well, otherwise GMail will be unable to correctly understand your headers. However, this effectively puts your server on the hook for any email forwarded; this is why I don't think you want to go there without also putting some kind of spam filter in place, otherwise I assume your server's reputation will deteriorate.
And is there a way to (manually) double-check those?
We desperately need a service that helps filtering email -- like, for example, something simple that just accepts reports about some email being spam or not, and creates a list of spam addresses.
This is not entirely true. Spamassassin, dspam, and all the bayesian based ones need constrant training and feedback loop to work. The time trainings lasts for is getting shorter and shorter, but it's not entirely inefficient. ( I'm running dspam on my mail box. )
Combine this with weighted blacklists ( postscreen in my case ), add dkim and dmarc checks and it's fine for a small provider. Far from ideal, but working.
I'd love to see an open source implementation of what google was referring to as domain based trust, but it's also a nasty thing.
I've recently tried to change my mail address from a .eu domain to a .net, and most of my mail landed in the recipients' spam folder. It's a fresh domain, no one ever sent anything from it, so I'd assume fresh domains are untrusted by default, which is really bad and is generally wrong. If the trust is not OK by default, spam is users consider it spam, that's going to kill domain based mailing, which is horrible, and gmail will be the one to blaim when we have not alternatives to few providers.
Question: Is there a curve where pushing this processing back to the phone will become possible? The most powerful counterpoint at the moment is battery life, but I do see that improving to a plausible point where this sort of continual processing is feasible?
While I understand that _some_ residential ISPs don't let you run services on your connection, policies like this make me sad because it means the web is becoming more-and-more something you need other people to do for you.
smtpd_helo_restrictions = permit_mynetworks,
reject_invalid_helo_hostname,
permit
smtpd_recipient_restrictions = permit_mynetworks,
permit_sasl_authenticated,
reject_invalid_hostname,
reject_non_fqdn_recipient,
reject_unknown_recipient_domain,
reject_unauth_pipelining,
reject_unauth_destination,
check_policy_service unix:private/policy-spf,
check_client_access pcre:${config_directory}/dspam_filter_access,
permit
This is in my postfix config. The invalid hostname and non_fqdn tests are working quite well against residential hosts: they don't have valid reverse DNS or an fqdn, and so they get eliminated fast.Another effort by anti-spam people to encourage network neutrality violations is the SUBMISSION port, port 587 as opposed to port 25, for the authenticated submission of mail by MUAs. As far as I can tell, the only possible utility of this split is to make it more politically feasible for ISPs to block port 25 out, i.e. to facilitate network neutrality violations. As such I do not view the SUBMISSION port RFC well.
(*Some people argue that PBL networks have no business sending mail because they have dynamic IPs. This is false; there are PBL-listed networks which assign static IPs, and it is at any rate a moot point. No specification requires that MTAs for domain outgoing mail and MXs (that is, MTAs for incoming mail) be one and the same, and there is only a technical need for the latter to have fixed IPs.)
See also my site https://violations.devever.net/w/Violation_Type:Rejects_Vali...
Personalized bulk mail (e.g. bulk mail with a small personalized offer) is an issue here. For that, partial encryption seems like a nice solution. As the template remains plaintext, reputation can include the template. As such, you could levy a lower 'tax' on such email as opposed to fully confidential email.
Disclaimer: if you haven't seen ads in the Gmail android app yet, please consider yourself lucky and kindly buzz off from this comment.
The reason why spam is an issue with email is, anyone can send emails and it's not always possible to identify the sender. Once encryption is deployed then the sender is associated with a public key and it's possible to establish a web of trust. Gmail could also manage identity-based scores based on all it's user's trusted connections.
I challenge this. In 2016 I think most companies don't need to rely on massive bulk mails anymore.
But then I always hated all forms of marketing...
> Botnets appeared as a way to get around RBLs, and in response spam fighters mapped out the internet to create a "policy block list" - ranges of IPs that were assigned to residential connections and thus should not be sending any email at all.
Your residential connection might not, but mine does.