Toshi: An Elasticsearch competitor written in Rust
github.com
github.com
> Toshi is a three year old Shiba Inu. He is a very good boy and is the official mascot of this project. Toshi personally reviews all code before it is commited to this repository and is dedicated to only accepting the highest quality contributions from his human. He will though accept treats for easier code reviews.
I've been tempted to just ship straight to a dB and skip all these crazy shippers and parsers and all the other middle men in the equation.
Also, why has no product unified monitoring and logging? AFAIK that's why Splunk is worth it is you have the budget (I dont)
I haven't checked, but I bet it's cheaper than ES for data that's mostly cold, like logs. You will need a separate monitoring solution though.
Perhaps I misunderstand your situation, but I don't see any "CREATE INDEX" available in Clickhouse, and thus won't "SELECT * FROM logs WHERE match(message, '(?i)error.*database')" require a full column-scan (including, as you mentioned, decompressing it)? Versus the very idea of an indexer like ES is "give me all documents that have the token 'ERROR' and the token 'database'" which need not tablescan anything
I only learned about the project 9 minutes ago so any experiences you can share about the actual performance of those queries would be enlightening -- maybe it's so fast that my concern isn't relevant
If your query is linearly scalable conceptually, Clickhouse is also linearly scalable. Per core performance is also pretty good. (tens of millions of rows per second on good hardware and simple queries, like most log aggregation queries are)
If you value being able to store arbitrary log files, ClickHouse is not for you. If you want to build your system to generate tables on the fly- ClickHouse might work.
Disclaimer: I work for Altinity which is commercializing ClickHouse.
In my experience SumoLogic is excellent for this.
Namely InfluxDB and friends
[1] https://jumptrading.github.io/influxdb/metrics/monitoring/de...
What is the difference?
Can you just extract metrics from logs via ES. Logstash even has a prebuilt JAVASTACKTRACEPART for java exceptions.
Metrics can always be generated from logs, especially structured payloads.
The docs have more info about both: https://www.elastic.co/guide/en/infrastructure/guide/current...
Disclaimer: I work on Kibana at Elastic
> I've been tempted to just ship straight to a dB and skip all these crazy shippers and parsers and all the other middle men in the equation.
You need the parsers. You want to find a needle in a haystack? Good luck without having things broken down into proper fields with metadata.
You need the shippers. Elastic Beats has full backpressure support so that when your cluster is busy, it can intelligently back off. Otherwise, you'll drop logs, or overwhelm the system to the point of uselessness....
> why has no product unified monitoring and logging?
Metricbeat from Elasticsearch + Grafana aimed against Elasticsearch to get you better dashboards and alerting.
Please don't reinvent the wheel on this one. Deploying ELK + Beats + Grafana is not that hard, there's tons of documentation, and it is a very stable product.
It was built for .NET devs using Serilog but can accept structured logs/json over HTTP.
Something similar I'm really hoping to see is Tantivy in a Postgres extension, so I can stop playing the game of trying to keep my search engine and database in sync. Seeing pg-extend-rs (https://github.com/bluejekyll/pg-extend-rs) on HN the other week got me thinking about it again. Does anyone know whether this is feasible or if anyone is working on something in this vein?
For my projects thats all I've been doing, in order to avoid the overhead of ensuring ES or Solr are synced with db.
[1] http://rachbelaid.com/postgres-full-text-search-is-good-enou...
You can convert existing dictionaries available to format Postgres understand, but this is annoying pain point if you happen to be an open source project like CMS or communication platform.
Apache Solr is more suited to search. Lots of document filters, query filters, the index itself is highly configurable and the ability to sort on multiple parameters is great. LTR is also something too good to miss out on.
1. https://www.elastic.co/guide/en/elasticsearch/reference/curr...
Lot of open source projects would benefit, if they had dedicated sales teams contacting firms day in, day out talking about their features.
People who run IT departments in the Enterprise world or even small firms lacking resources to keep up, just pick tools and make software decisions based on who reaches out to them.
Invest in a sales team and you penetrate markets that don't spend time monitoring developments in the tech world which is really the majority of all orgs. Elastic has done that quite well and is reaping the rewards.
Solr was there earlier. In many ways, Elasticsearch was a response to stuff Solr did not do, or did not do well. Like clustering for example. Solr has of course added this since then.
I had a brief play with it for a project at work and it was super straightforward to get running.
I found this part of your comment interesting, given that Rust is in someways being offered as an alternative to C++ for similar use cases.
What would you like to see different in a language that you see as the common issue with C++ and Rust?
edit: I misread the parent comment, the question should be disregarded.
I would also be very interested in that. Elasticsearch usually works very well for me, especially the latest version, but it feels very heavy, and seems to only be getting worse in that regard.
I should probably work on a proper scenario for schema evolution.
I actually strongly considered that just with Solr, which has the extreme benefit of using the same query language under the hood, but the more I scratched the more I found it would be a horrific amount of work
At least in my experience the web interface is a huge gap. Kibana is ok but Splunk and it's query language and visuals are much better. Anything that competes with Splunk is great imo.
Subscribe to the blog (we are non-spammy) to get the email announcement.
> https://github.com/toshi-search/Toshi/blob/master/roadmap.md
It seems that the main added functionality of ES, namely clustering, will still take a while to be implemented.
At the same time, I’m curious on the performance differences with postgres. I’ve been able to get very quick queries from Postgres:
https://austingwalters.com/fast-full-text-search-in-postgres...
The only time it’s slower to the point of using elasticsearch (for my usecases) is something like log search (so far).
But there's an attitude attached which amuses me:
> Toshi will always target stable Rust and will try our best to never make any use of unsafe Rust. While underlying libraries may make some use of unsafe, Toshi will make a concerted effort to vet these libraries in an effort to be completely free of unsafe Rust usage. The reason I chose this was because I felt that for this to actually become an attractive option for people to consider it would have to have be safe, stable and consistent. This was why stable Rust was chosen because of the guarantees and safety it provides.
It's an admirable goal, though the fact that it's stated prominently as one of the first things on the front page gives off somewhat of a "doth protest too much" vibe. It's not like "safety" is rare these days. Any project written in Go, Python, Java, C#, Erlang, JS, and a myriad of others will be "safe" as far as memory access is concerned, and in many cases this safety will be easier to achieve than in Rust. As far as error handling safety, so far exceptions seem to be more expressive, though the jury is out there.
Basically, if a project stays away from C and C++ and libraries written in them, it's more likely it will be hit by a hardware problem than an inherent language safety / security issue. Luckily, "safety" is for the largest part the default for modern projects.
So, you're writing in Rust and you'd like to write code which has the absolute most aggressive performance possible. In Rust parlance this _maybe_ means you need to do unsafe things: turn off bounds checks, fiddle with the raw memory of a structure, allow multiple threads to access the same memory without synchronization, interact with mutable globals and so on. That's great, but, you've potentially opened up your program to crashes or security issues. If you're a solo author that's a trade-off you can make based on your needs. Here's the rub: if you use similar techniques in a crate -- a shared library for all to use -- then you've opted everyone using your crate into the same trade-off, one which they might not have otherwise chosen for themselves. What the Toshi project is saying here is that it's design preference is to avoid opting into this trade-off, preserving all the guarantees that Rust can provide at _possibly_ the expense of absolute performance.
There's a safety-focused subset of the Rust community that takes the presence of an 'unsafe' in a body of Rust code very seriously and this project is participating in that conversation.
This allows the “unsafe” libraries to have fewer lines of code and more isolated testing, with greater coverage.
I'm still thinking about the rest of concurrency problems and side channels that Rust's type system doesn't cover. Trying to find out what's as easy as above. For concurrency, I'm eyeballing Eiffel's SCOOP, DTHREADS, and eventually will study whatever Pony is doing.
Not sure I totally understand (or if I do, totally agree). For all FFI in Rust, we have to drop down to unsafe interfaces. I like the model that's generally happening in this area where there is an auto-generated sys crate (with bindgen), then an FFI crate that does all the Rust <-> C interop. This tends to work pretty well.
A lot of C/C++ (native library) validation tools just work with Rust artifacts, so I don't personally see a lot of value for writing in C as well as Rust, unless we're talking about a rewrite from C to Rust.
> still thinking about the rest of concurrency problems and side channels that Rust's type system doesn't cover.
What things are you thinking about beyond the Send/Sync traits? I've found those to be very expressive, and appropriately restrictive.
Tools like RV-Match and Astree Analyzer can prove absence of entire classes of errors using static analysis. Frama-C and SPARK Ada can do that with annotations with high amount of proof automated. There's optional, runtime checks for stuff not proven if you don't want to or can't do it by hand. C also has lots of open-source tools for static/dynamic analysis, test generation in many forms, and so on. In my Brute-Force Assurance concept, you convert a program into C and/or Java to throw all the automated tooling you can at them, fixing whatever real issues are found. Then, the last benefit is C has a formally-verified compiler should someone want to eliminate compilers as an error source. Rust-to-C is still valuable for all these reasons.
So, is the Rust tooling at the level that you can do all that without converting it to C?
"What things are you thinking about beyond the Send/Sync traits? I've found those to be very expressive, and appropriately restrictive. "
I don't use Rust yet. I just know what I learn from folks like you who do. When I studied it, the docs said their type system blocked some but not all concurrency problems. I don't know where it's at currently on various types of races, deadlocks, and livelocks. Those are the main problems improved models or analyzers should try to solve.
But comparing Rust to C is difficult. On the one hand they compile down to the same thing, on the other when not using unsafe, the type system itself allows you to express strong proofs, especially with state machines.
I don’t generally need to make use of these tools, I would point you to the ring project, where they are very interested in formal proofs: https://github.com/briansmith/ring they might have some interesting options and a few of the people over there are very capable in answering these questions.
> I don't know where it's at currently on various types of races, deadlocks, and livelocks.
For dataraces, there is a very strong story. In terms of deadlocks, I’m not aware of anything here. For livelocks, I think it’s generally possible define the state of a system such that you can make sure you aren’t in conflict with other threads, so better, but not fundamentally different than other threaded languages. In other words, if you define cross thread state in an appropriate way, you can prove that you can’t get into a livelock situation.
I was in that comment section bringing up RV-Match. The difference is that RV-Match is a bunch of static-analysis functionality built on a comprehensive semantics for C. KRust is a tiny subset of Rust with RV-Match-style analysis. I did bookmark it in case it could be useful for someone trying to build that.
"where they are very interested in formal proofs"
The only thing I mentioned that would be doing a formal proof was the lightweight stuff like Frama-C and SPARK Ada that don't require proof so much as annotation (eg like borrow checker) that run through automated provers. The rest was all push-button or no-proof-needed tools that say something has no errors, specific errors, or a mix of them and false positives. RV-Match and Astree Analyzer will straight-up tell you that specific errors don't exist with low, false positives. The test generators that work on structure of your code get deep into lots of errors from different combinations of inputs and control flow. None of these require proof. People building these usually test them on FOSS projects, sometimes the same ones, often finding new errors in them.
"In terms of deadlocks, I’m not aware of anything here. "
Good to know. That be the focus area for now since you said livelocks are a good situation. I'll still keep an eye on the side for anything checking that.
"they might have some interesting options and a few of the people over there are very capable in answering these questions."
Thank ya very much. I'll hit them up, too.
But few of them has as strong guarantees about your safety as rust. Python, Erlang, JS will not enforce anything and raise recoverable errors on mismatch. Java and C# will enforce a bit of correctness and raise recoverable errors. Go is the most strict, but happy to panic at runtime if you cast to the wrong interface. And that's before we talk about things like defined integer overflow behaviour, safety without concurrent execution, etc.
Those languages do protect from accidental memory overflows, but rust offers a lot more in safe code.
This was in response to the point that Go panics, and I didn't think it was perfectly fair to use that as an example (I have others that I would use) since Rust also panics in similar situations.
> as well as some numerical operations (in debug mode) that will panic on overflow
That's a good thing. You normally want to find them in debug mode.
> Indexing into arrays/slices outside bounds will also panic.
Kind of. There's .get which returns an Option to do it safely, but you're right that basic indexing can fail.
There are situations where panicking is something you want to try and guard against, so I hope something like this gains traction and is able to be used more generally: https://crates.io/crates/no-panic