Sieve: An Email Filtering Language (RFC 5228)
rfc-editor.org
rfc-editor.org
I've been leaning towards local processing on an mbox folder... that would also let me build up multiple Bayesian filters to try categorizing my other mail - why does almost everyone only auto-categorize spam? Google kinda broke out of that with their "social" and "important" semi-categories, but it'd be pretty easy to build in local clients, yet I've never seen it.
For bayesian junk filtering to work, you train with messages marked either as junk or nonjunk. I suppose you could train based on positive tag, and assume the message has the opposite signal if it didn't get that tag. My historic email is only classified for (non)junk. If I would start more classifications, I would have to ignore the existing messages.
Fyi, I'm working on https://github.com/mjl-/mox, which also includes a bayesian filter. This sounds implementable.
All reasonably identifiable by word choices, and often at extremely different urgency / priority levels.
But I would personally also have an "interesting" category for newsletters and whatnot that are more likely to get a full read (newsblur has a weak version of this for my RSS, it's nice). And an "urgent" category that notifies more visibly than others. "Recruiters" is also periodically useful - a bunch of cold emails filled with buzzwords, should be identifiable a mile away.
But underneath it all it's mostly that it confuses me that there's so little ability to experiment here. Shared always-updating spam filters makes sense, but there's a lot of similar things people could try unique to their needs that are just totally unsupported. I don't want better global categorization (gmail), I want personalization. Email is weirdly non-personal in the vast majority of systems, and I suspect it's part of why people burn out on it and switch to a never-ending churn of apps that are a better fit for one piece or another.
https://www.fastmail.help/hc/en-us/articles/360060591373-How...
– A notorious spammer that has been spamming me for over a decade that the Fastmail spam filter does not trap[*]. They are easy to screen out via Sieve rules using the "From" header. Except, for 1 out 5 spam emails from them the Sieve rules do not fire, and the spam gets into the inbox.
– Simple email filing into different folders using From/Subject headers. It mostly works except for just a handful senders whose emails Fastmail Sieve do not fire upon no matter how I modify them. They just. do. not. work. I am talking about the most basic From/Subject filing rules.
[*] Spam filtering in Fastmail has been another pet peeve of mine. It is average at best, and mediocre on average. Some spam emails that slip through need to be sent into the spam folder (that has the training enabled) for more than 10x times before*
It automates a lot of manual cruft and I’m endlessly thankful for it.
I do wish it was more widely supported, and I try to do my part to encourage adoption (I maintain several sieve related packages for Arch Linux).
---
Between the insufficiencies of IMAP for server-side use and the disparities between the "winners" of the space, the lowest common denominator of what's supported for email is abysmally small.
I'm not sure Sieve is a language that enables such usage. I know maybe it's a domain-specific language, but since it's a language capable of programming the server, I would rather expect more from it than "it's just yet another way to configure your mail filter".
What you can do though is edit your rule as a seive filter and just add multiple rules in there and it works fine.
believe it or not, but many hugely popular applications from 20+ years ago didn't even implement indexing. in this case if you had too much mail your imap client would just time out.
well through the mid-2000s you even had to use a third party plugin for outlook called "lookout" if you wanted searches to not take minutes.
Folders (mailboxes in proper IMAP lingo) had hand-built indexes. Good stuff. Credit to jgm, the original author.
i also think it was the first support for sasl(?) encryption upgrades for legacy text/tcp mail protocols
also, fun sidebar: indices vs. indexes, both are apparently valid english... but it seems computer people have adopted the latter almost exclusively. never noticed it before...
Kind of funny that SASL is the most durable piece of the effort.
(I doubt this is entirely accurate. I wasn’t there for a lot of it.)