102 karma · joined November 18, 2010
It also serves as a nice showcase what kind of tools can be built with using nothing but the public APIs
I'm one of the maintainers of tremor, happy to get together and talk about rust event processing if you ever want to :)
Or feel free to shoot dm on twitter (@project_fifo) w/ your e-mail and I will make sure you get an invite :)
Storing labels in a row based system (like SQL) allows querying by value, not column name which takes advantage of all optimizations and indexes making it a lot faster.
That said there is nothing forbidding someone to do both, DalmatinerDB, for example, uses a column-based format for metric values but a row-based format (PostgreSQL) for dimensions.
There are some more subtle differences like the for BSD Linux emulation is a global setting and 'lx jails' are just jails w/ a Linux userland while branded zones are special kind of zones.
The only thing to criticise (to a degree) would be that if the manual isn't followed precisely during the installation you can end up with a busted setup.
For why the choice. It's a solid distributed system, the concepts it uses the same principles as FiFo (masterless setup for high availability). Being written in Erlang means it works flawlessly on SmartOS and FreeBSD plus if you already use erlang gives the advantage to be able to look at the code if needed.
The LeoFS team is very quick to respond, works extremely diligent and takes their work serious (which is a big plus).
Even on the test system which gets brutally shut down (aka plugs pulled) about once or twice a week the installation works flawlessly even after a few month of this torment.
Of cause as always YMMV ;)
They have however been excellent open source citizens and contributed back improvements, bug reports and suggestions.
But I'm kind of curious, are you using Dalmatiner directly? And if so what is your use case?
UDP is no longer used (this is outdated sorry for that), the connection is TCP now as it turned out over all the performance was better.
Dataloss can still occur since DalmatinerDB keeps a cache (which other metric stores might also do). A lot of that can be mitigated by using N=2 (or 3) to store data in multiple nodes that will reduce the chance of dataloss significantly. Keeping caches isn't uncommon however, to ensure full consistency it requires a kind of transaction from that goes from client to server to client, which is 'really' costly and I am convinced not worth it for metrics, a few seconds of lost metrics doesn't warrant the cost of that.
What DalmatinerDB was build of is handling high numbers of metrics like CPU usage, memory usage, etc, everything where for every point in time you have exactly 1 value. In that space it can handle millions of metrics (read millions of inserts) at the same time where a system like postgres would stat to have problems in my experience.
What they do is like saying "Hey I want to buy your house, what you don't want to sell it? How about you make it cheaper then so I can afford it!"