Replacing Elasticsearch with Rust and SQLite (2017)
nickb.dev
nickb.dev
I've written a few tools to help work with FTS in SQLite:
- https://sqlite-utils.datasette.io/en/stable/cli.html#configu... - command-line utilities for enabling full-text search against an existing SQLite table
- https://sqlite-utils.datasette.io/en/stable/python-api.html#... - that same functionality as a Python library API
- https://docs.datasette.io/en/stable/full_text_search.html - Datasette spots full-text search enabled tables and adds a search interface to them
I also put together this article exploring classic search relevance algorithms with SQLite - SQLite FTS5 has relevance built in, but FTS4 leaves it as an exercise for the developer which makes it a really fun tool for understanding how algorithms like BM25 actually work: https://simonwillison.net/2019/Jan/7/exploring-search-releva...
No support yet for regexp searches over trigram indexes, though, unfortunately (Postgres has this.) I personally think that much more expressive indexed search syntax (more so than just regexp, even) is going to be the future and am actively working on a search engine in Zig for this purpose.
[0] https://sqlite.org/fts5.html#the_experimental_trigram_tokeni...
https://devlog.hexops.com/2021/postgres-regex-search-over-10...
e.g. I discuss the fact that pg_tgrm indexes being built does not take advantage of multiple cores, so at large scales you'll benefit from splitting data into multiple tables.
Doesn't support languages other than English. At least, that was the case when I last tried it.
I tell people Xapian is the SQLite of search.
I've perused their docs and the immediate question for me is, why would I use this over another more popular search library or database that already has a wealth of stackoverflow and blog posts?
I'm sure it's a fine tool, but I don't see a reason why an enterprise should hedge their bets on this unknown.
[1]: <https://notmuchmail.org/>
Again, I don't think Xapian is a bad tool or that it's underperforming by any stretch, but OP should not be surprised at all that most people have not heard of it.
I recommend checking out scout[0], which, I think, can be a good replacement for Elasticsearch in some cases. I'm also working on an Elasticsearch replacement built on top of SQLite for my litements[1] project, but it will still take a few weeks to have a working version.
[0] https://github.com/coleifer/scout [1] https://github.com/litements/
I appreciate - and often indulge - in writing stacks of software that are better implemented other open source, but I wouldn't say that the article describes a solution that is viable as a replacement for Elasticsearch.
Edit - i.e. where ES is overkill for the use case, however for an appropriate use case ElasticSearch is amazing
1. No network round trips 2. SQLite runs in the same process as your app, so reading from SQLite is basically like calling a function in your program (from a performance point of view).
> SQLite runs in the same process as your app, so reading from SQLite is basically like calling a function in your program
But SQLlite needs to write to disk, which will be much costly.
[0] - http://www.grantjenks.com/docs/diskcache/cache-benchmarks.ht...
Would it be conceivable to replace Algolia in this way for static data sets (i.e. those that don’t need the write/update API)?
I feel like most of the time, SQL DSLs end up being bad. SQL is already a very high level language. There are a ton of examples of how to do stuff with SQL queries. DSLs, even if they have a way to express the SQL, are not nearly as widespread as SQL and then you will have to spend time figuring out how to translate the SQL into the DSL.
In addition, with libraries like sqlx for Rust, you can also get the type safety of a DSL using regular SQL.
I would say that as a developer, the time spent getting familiar with SQL is a good investment that will likely pay off across many projects and programming languages.
This is a really excellent exercise for understanding material need for refactoring.
I don't think I share the author's conclusions but it makes a really good test case.
"At idle, Elasticsearch uses:
1GB of RAM (out of 8GB total)
30% utilization of a CPU
1GB of Disk space after a week
This doesn’t sound bad, but considering this server houses a dozen containers and the next CPU intensive container uses only 2% of a CPU – Elasticsearch is too heavy. For a server that gets 15 requests a minute, the resource consumption of Elasticsearch doesn’t justify it’s use. What would happen if I received a spike of traffic? I feel like the server would fall over not because the app couldn’t handle it, but log management couldn’t handle the load."The thing is 'uses more resources than the thing next to it' is maybe a nice referential starting point, but it doesn't mean much.
What is the actual cost/risk/benefit work out to be? Cost in terms of development, maintenance, hosting, opportunity for future expansion?
Are we going to need to 'do more soon'? Are the hosting costs even material? Can Rust be easily supported?
The bit about 'what if there are more requests' does give pause for thought, but, it could be that there's a threshold/minimum that the service needs to operate, above which there's only marginal, incremental resource consumption. I don't know, I'm only pointing out the possibility.
And the choice of Rust, the author denoted has a few nice attributes ... but is this a personal choice ... or an optimal choice? Would the Python or Java solution be cleaner? Elastisearch is based on Lucene (Java), so it might be possible to do something very fast, powerful and extensible there as well.
Thanks to the author for both 'doing it' and 'writing about it' - but I think the meta issue here hinges is the technical product case.
Right below where you quoted the author he stated his biggest risk - denial of service not due to his site having performance issues but his logging solution not being able to keep up if he had an unanticipated traffic spike.
That would indeed be a rather nasty case of self-ownage; the benefits of avoiding that should be pretty obvious.
https://appliedmachinelearning.blog/2018/07/31/developing-a-...
?
"""
Greylog Extended Log Format (GELF) is a log format that avoids the shortcomings of classic plain syslog:
* Limited to length of 1024 bytes – Not much space for payloads like backtraces
* No data types in structured syslog. You don’t know what is a number and what is a string.
* The RFCs are strict enough but there are so many syslog dialects out there that you cannot possibly parse all of them.
* No compression
""" -- https://docs.graylog.org/en/4.0/pages/gelf.html
It can be sent over UDP or TCP. Looks like each log entry is a JSON object, with standardised field names/formats.
Interesting how we can refine use-cases and come up with simpler solutions that work as well or better. It's not a general replacement for Elasticsearch, notably may not scale for searching all the nginx logs for a distributed multi-tenant system.
My gut feeling says it won't make too much of a difference.