FastMail's Email Search Architecture
blog.fastmail.com
blog.fastmail.com
I use it from the emacs notmuch mode.
http://www.djcbsoftware.nl/code/mu/ http://www.djcbsoftware.nl/code/mu/mu4e.html
But I don't use emacs, and that seems to lead to subpar support and crazy hacks to get something up and working, unfortunately.
Me neither, but mutt-kz had built-in support for notmuch. Not the hacky kind that calls 'notmuch', but it actually links against libnotmuch.
I came from regular mutt so there was no learning curve, but that shouldn't be too bad either once you get used to those keybindings.
Once the indexes are up to date, we can switch back to being masters again. We index on all the replicas independently so that they are always ready.
With sphinx, we found we had to start and stop daemons all over the place to manage memory, and it was just unworkable. It was either that or run one big index per machine, but there are operational reasons I'd rather not be doing that. We try to keep everything user-sized.
That said, there's still stub Sphinx code in there. Both engines are have GPL licensing on them, which means compiling against Cyrus (BSD licensed) causes a non-BSD licensed end result. Not an issue for us, since we publish all our Cyrus code anyway.
There is talk of building an Elasticsearch backed into Cyrus as well - feel free, it's all open source. We'd definitely take the patch if it's good code (he says with his Cyrus Project Board Member hat on rather than his FastMail Director hat on)
Solr (well, Lucene) has awesome natural language stemming abilities, many more languages supported than Xapian. In particular it's much smarter than Xapian about Chinese. But a) the memory requirement makes running multiple shards on the same machine hard, and b) nobody in the company wanted to learn how to handle Java operationally.
EDIT: since 2012 Sphinx appear to have made a public mirror of their internal tree at https://code.google.com/p/sphinxsearch/source/checkout
[0] - http://lucidworks.com/blog/podcast-solr-at-scale-at-aol/
Not saying you can't do it with Solr or that Solr doesn't scale, it does. You'll just have an easier and more fun time doing it with ES.
Couple of related/examples:
http://highscalability.com/blog/2014/1/6/how-hipchat-stores-...
The reason is that you want to keep the inverted indexes sorted on disk, but you don't want to sort the entire index every time you update. So you create one mini-index per update and merge them lazily when you get too many of them.
We use Elasticsearch elsewhere (ELK stack), but not for mail search.