I guess you don't see much demand because for a lot of use cases the basic setups are good enough.
I guess you don't see much demand because for a lot of use cases the basic setups are good enough.
I ended up trying to reinvent Solr for the client, realizing after about two days of trying to reinvent stemming and indexing, that this was stupid to do on someone else's time, and called the client to tell them that I'm moving to Solr, and I got the project done before-schedule as a result.
====
I think for 99% of usecases (involving search), Lucene/Solr/ES is perfectly fine. However, I do absolutely hate that some companies have decided to make it their primary database.
EDIT: I just want to make it clear, I think it's totally valid to try and reinvent Solr for fun, or if that's something you're paid specifically to do; nothing is perfect, and I am actually a big fan of the "if it works, break it and make it better!" mentality.
TLDR: Lucene < 7.5 won't merge segments larger than 5GB (default) unless they accumulate 50% deletions.
Delivering a conference talk [1] later this year about it.
[0]: https://www.eivindarvesen.com/blog/2018/09/16/elasticsearch-...
[1]: https://2019.javazone.no/program/3f7cd8a7-a9ea-4874-a7dd-531...
I don't think the GDPR regulatory agencies are operating at a technical level that they would make an argument that it was not a good enough deletion.
Finally I have to ask this part: assuming ES is not your primary database, how does this get around the GDPR issues? If someone wants their data erased you are supposed to erase it from wherever you store data, I suppose this means ES when it indexes a primary store and finds it has deleted data actually deletes it but if it is told to delete something it keeps it around?
If it is easy to delete specific users from the primary database, the deleted users will naturally disappear during the next ES reindex.
Edit: The old index is deleted at the file system level.
If the reindexing occurs daily or weekly, perhaps this will satisfy GDPR.
There are other good reason to not use ES as the primary data store. First, it isn’t entirely reliable. It’s good and I’ve never seen a corruption, but ES and Lucene’s history isn’t as a reliable database. Second, if you want to change how you index, it is a bit easier to do if the source data is outside of ES.
How is GDPR compliance of having data in Elastic Search influenced by it being primary vs. secondary storage?
Basically: you reindex ES periodically, so when a user is deleted from the primary, it will disappear from ES upon the next reindex. The old index is deleted at the file system level.
at what point would you be able to 'delete' data without being in violation of GDPR?
Though the EU has said it will consider intention etc. there's really no way of knowing for certain until when and if it's settled in a court case.
Fix it 'til it breaks is what I always say :D
Needless to say my 3D printer had a lot of down-time haha.
I inherited a legacy Mongo solution, and all the data is duplicated and indexed in ES, so I've always wondered why we're using both. Mongo has none of the SQL capabilities that would make my life easier, and the types of queries allowed by Mongo could be done with ES.
What are the negatives of ES alone?
The v7 upgrade to a new cluster protocol (zen2) has improved things but overall the system has a long history of losing or destroying data. It's better to have a primary OLTP system that's ACID and reliable while using ES as the secondary search source. You can also remove the _source field if you just need matches without the original content.
It's common to see pattern used with a relational database since, as you can see, Mongo doesn't buy you much else as another document-store.
Thing is, there's heavy demand for something more performant than Elasticsearch, so eventually the market will provide.
Meanwhile, Redis Enterprise is trying to grab some 'share with RediSearch, which has some severe caveats IMO that make it not a great fit for most.
https://github.com/tantivy-search/tantivy
That said - it's effectively Lucene rewritten in rust, so the main win is some performance gains. Lucene has spent a ton of time getting the details right, and it's unlikely we'll see an order of magnitude of innovation in that particular space. At the higher level for querying / query understanding it feels like there's still more technological room to grow vs the lower level details.
It is not exactly a port but yeah. tantivy is strongly inspired from Lucene.
> Lucene has spent a ton of time getting the details right, and it's unlikely we'll see an order of magnitude of innovation in that particular space.
Have you checked out the perf gain in Lucene 8.0 ? Block-WAND proved you wrong.
Tantivy is a cool project, but I have to say the part I love most about it is your blog posts on it. They're a great introduction for people who are unfamiliar with the underlying tech of search engines.
Thanks a lot! I am not a native speaker, and I often feel very bad at conveying engineering concepts. The positive feedback is actually very helpful :)
Disclaimer: I’m a previous employee but have no economic interests in this as it’s not a publicly traded company.
Vespa [1] would like to have a word with you.