Elasticsearch for Beginners: Indexing Your GMail Inbox
github.com
github.com
At least it's a git repo so you should have local copies, but still, that caveat applies.
- https://github.com/elasticsearch/elasticsearch-py (low level lib, from ES)
- https://github.com/elasticsearch/elasticsearch-dsl-py (high level lib, from ES)
- https://github.com/mozilla/elasticutils (high level lib from Mozilla)
There are a few more, but they are either obsolete or don't have much traction. There's also django-haystack, but that's specific to django.
You can try it out here: https://www2.hsnet.nsw.gov.au
Try searching for something like: psychiatrists near sydney cbd
The site is self is mostly just a front-end (as designed by / for a client) running across several Docker containers for the database (PostgreSQL) which is indexed into Elasticsearch and queried via the Elasticsearch API / Python ES Libraries.
If you're interested in Elasticsearch with Python / Django check out our (pretty crappy at the moment) tech blog: https://ixa.io or our github: https://github.com/infoxchange
I realize this tutorial is just meant to get started with elasticsearch and not meant as a tool to make your email searchable. Still would be interesting to take this to the next level.
I don't know how to invoke the script properly...
I've tried so many ways. This seems like it would give results... though it does nothing much.
python index_emails.py test.mbox
Any help or tips are appreciated! This has been a fun project so far. Stumbling at the end. Thanks!
python2.7 index_emails.py --infile=test.mbox
above is working
There are a number of authentication solutions, and they will require additional configuration -plugins like jetty and elasticsearch-http-basic.
If there's a demand for this, it might be worthwhile to build IMAP servers with more indexing. It's easy to request searches with IMAP, but the performance can be a problem for IMAP servers that aren't real databases.
I don't think this has much to do with the search in gmail not being sufficient or broken.
Since it's implemented as a "library" of sorts, there are interfaces for emacs, command line, GTK, mutt, ...
Anyway if you do feel like you want to accomplish the stated purpose of finding which emails are taking up space, you can search in gmail with the word "larger", as in "larger:20MB".
SELECT id FROM pages WHERE title LIKE "%elastic"This means that at best it will need to do a index-scan (in practice a full table-scan on the index consisting of a subset of the table-data). The results of a query like that can't be obtained fast, at least not milliseconds fast, and its not well suited for caching or reusability.
There are (obviously) ways around this, but they all involve indirection and workarounds.
Anyway, as a rule of thumb: Any time you want to do general full-text search queries on relational data, the LIKE operator will very quickly meet its limits and depending on how much outside its bounds you want to go, you will find yourself having to do quite a bit of work.
And in those cases you mihgt want to use the full-text capabilities of your RDBMS if it has any. Or maybe something like Elasticsearch might be a better fit.
Disclaimer: I've never used Elasticsearch myself.
Think of search engines like ElasticSearch and Solr as being purpose built for "search" rather than ad hoc querying.
They offer more advanced searching features like faceting and synonyms, if your example had been "SELECT id FROM pages WHERE title LIKE '%dog'" you could set things up so that matches for 'dog', 'dogs', 'doggie', 'puppy', 'pup', 'canine', and 'mans best friend' all returned the same results.
If you under a million rows, LIKE is reasonably performant.