ElastiQuill: A modern blog engine built on top of Elasticsearch
github.com
github.com
While it is possible to use ES as your primary store, doesn’t mean it is a great use-case. Pushing ES to durability takes resources (people, time, compute) and rigor.
For a blog site this feels like overkill. Your priorities are durable low-freq writes, and more freq-reads. Hands-free day-to-day maintenance, and simple recovery mode such as restore. Caching (and optionally queuing) can absolve sins of many topologies to fulfill these requirements.
ES brings a much larger surface area than just store blobs and metadata. Big tax to gain free-text search - which while interesting most users will almost certainly never use (thank Google for that)
Focus on a cheaper, lighter, database engine that is durable, checkpoints regularly, and has simple, testable backup/restore processes. Ship the backups somewhere e.g. S3
* (They'll do it for you)
To your point, it does feel like "using a cannon to kill a mosquito". But I still had a better experience with it than even sqlite.
It reminds me of https://nodebb.org, the "forum built with Node.js" (though they've rebranded away from that) which always struck me as the opposite place to focus when selling software.
For example, none of my problems with Wordpress are related to its choice of database. Just like porting Wordpress to Elastic Search don't fix any of my issues with Wordpress. A blog is almost entirely defined by its UX.
On a better note, the platform seems pretty polished.
A blog is almost entirely defined by its content. I read a lot of run-of-the-mill average to crappy Wordpress or Blogger blogs that neither load fast nor look flashy (or sometimes even pleasant). I don’t give a shit because I read them for the content, mostly from an RSS reader.
However, I think as a public facing content engine it’s fine. I know of one major e-commerce brand that does this for products. This works great for their use case, given how easily Elasticsearch lets you filter, sort, search pretty seemlessly. But they can always rebuild Elasticsearch from another system.
Somewhere around 5 or 6 the tides turned. What was a big issue is how easy it was to overload the cluster and how data would be lost in that state.
To be safe you should have regular backups anyway. I wouldn't trust ES to handle very critical data which you wouldn't ever want to lose a bit of, but I have much less worry these days about the operational overhead of fixing an cluster because failure in experience is just much less likely.
It is still though, something which is less "set and forget" like many sql engines can be in less critical or intense workloads. You would want someone in your org to really know ES.
However, it's not a database. I've actually abused it as such and it's fine. You get optimistic locking but no real transactions. Search is eventually consistent unless you call _refresh (but you shouldn't), etc. Bearing in mind it is not a database, it is not intended to be used as a database, and probably will never be a database, it actually works fine as a database provided you don't do a lot of updates (write model is append only).
If safety is a big concern for you, obviously use something else. But for a blog it's completely fine.
If you use ES for what it's designed to be and make it a projection of your underlying durable data store, you can do things like rebuild your index or change your schema without fear of data loss.
Its a document store, not a database. MongoDB is too but they try to hide it by bolting on all the stuff traditional relational DB's have like ACID and transactions so you wont realize what a poor choice it is as a database when compared to say, Postgres.
It doesn't change the larger point that using ES as your primary database is not playing to its strengths. You're better served using a transactional data store and building your indexes from that.
While dramatic progress and improvements have been made over the past 9 years, sometimes things still go bad and indices get corrupted. When this happens, it's necessary to reindex the data. There are also additional situations where reindexing is required. So the safe advice is: Always have the authoritative data source elsewhere (in a reliable data store or database of some kind) and then load the data into Elasticsearch from there.
Postgres or even MySQL will be a safer bet when data integrity is key. Then it's only a matter of indexing the data from the DB into ES.
Using Elasticsearch for a blog is a fun you idea, but ultimately is likely to be overkill for a personal blog site.
I haven't seen indices getting corrupted in the last 4 years of having worked with it in production systems heavily. Not saying it can't happen but just haven't seen since elasticsearch 2.x
Meanwhile, if you want a small deployment that also does full text search with a bunch of "smarter" features and some analytics, it would make sense to only use a single data store.
All these in theory, as I haven't checked out the actual project.