InfluxDB – Open-source distributed time-series, events, and metrics database
influxdb.org
influxdb.org
I think there is a hackernews rule someplace that a more interesting tech alternative shows up right when you decided to go with something else.
For reference here is the list I created when researching these:
http://opentsdb.net/overview.html Built on HBASE
http://www.gocircuit.org/vena.html Built on go + go'circult uses google's leveldb
https://code.google.com/p/kairosdb/ A rewrite of opentsdb which can use Cassandra
http://blueflood.io/ Built by Rackspace, decent but still seems a bit immature
http://graphite.wikidot.com/ Obligatory Graphite reference (uses whisper, new backend called 'ceres' is being developed)
https://github.com/agoragames/kairos (yet another 'kairos', alternative backends for graphite - SQL, redis, or mongo)
Riak seems to be popular with SASS metric providers (hosted graphite, boundary). There isn’t any code but there are a couple of talks that explain how and why they went with Riak:
http://basho.com/hosted-graphite-uses-riak-to-store-all-cust... http://boundary.com/blog/tag/tech-talks/
Can you share anything of your experiences with kairosdb so far - what's the use case and how has it performed?
We looked at other options (including open tsdb and graphite) before building this.
Https://github.com/imvu-open/istatd
http://blog.apiaxle.com/post/storing-near-realtime-stats-in-...
I need to read up on rrdtool as well, but I wonder if it would make much difference (good or bad) to store the mean or other average as the "higher up value" (ie: the average of the past 60 seconds as the minute value) ?
It is, however, rather expensive.
kdb+, on the other hand, is a time series database that works perfectly well as a general database, with a query language that is at the same time infinitely simpler than SQL and yet much faster, more expressive and useful. There's a learning curve, it is steep, but it is well worth it.
I would simply use functions and operators over time series or data frame types. Perhaps take a look at the R zoo library for examples of more advanced things people do with time series.
I'm curious, did you find the SQL dialect readable and understandable?
But SQL is really a bad option in the moodern world, especially since (I estimate) about 90% of SQL statements are built programmatically; Thus, a query language that is easier to construct from code makes a lot more sense. (And not, that's not JSON - some form of algebraic notation or LISPish notation makes much more sense).
Also, SQL semantics are horrible if you have order involved (as you always do in time series).
These things (like SQL) make sense if you assume that the input is (a) written manually, and (b) by people who are not expected to do this "professionally". Neither is the case of HTML nor SQL anymore.
(Seriously, SQL was originally marketed to managers with the idea that "it's just plain english so you can do it yourself, and don't need programmers!". You know how well that worked out)
Pretty well, actually -- lots of nonprogrammer analysts use SQL for queries, and IME the ones that do consistently are better able to answer questions based on data than the ones that use "friendly" query tools, that inevitably end up being much more limited in practice, and requiring a lot more support from both programmers and DBAs to make the data that is already available accessible through.
Unfortunately, lots of environments prevent direct SQL access to DBs for "security" reasons (as if mutiuser DBMS's didn't have role based access controls as a core feature)
If you can properly do inner/outer/cross/asof joins to get to the data you want, the english-like syntax is just a burden - two queries that seem similar in their English more often than not produce completely different results because of SQL's 3-value logic, the way NULLs are joined, and various other things like that.
I don't think that's not true -- there's lots of adults that have anxiety around "maths-like" notations largely as a result of issues with maths education, cultural factors, etc., despite being able to intellectually handle the relevant manipulations -- and lots of those people end up in non-technical business positions that end up having to deal with data. Lots of the non-IT people I've seen using SQL definitely fall into that group, and I don't think they'd be as proficient with a more algebraic syntax like the comprehension syntaxes used in many modern programming languages.
Conversely, the people that can be proficient in those syntaxes almost certainly can be proficient in SQL, though they may complain about its verbosity.
Sure, in a perfect world where the cultural context was different, this wouldn't be necessary. But we don't live in that world.
How robust/scalable in your opinion the backend is at this stage? I'm just trying to set my expectations properly when checking it out.
Thanks, Sasha
The single node performance at this point for writes is tens of thousands of points per second if batched, and for reads we haven't optimized yet. Queries that only have to go through a few hundred thousand points should return in < 1s. Of course, those numbers will be highly variable depending on if you're writing a series with 1 column or 30 columns. We'll be adding things to make that better.
For now we're focused on creating a developer friendly API and building out clustering.
The interface I'd like would be something close to numpy (or matlab/r, if those are more familiar). Let me do vectorized operations and write functions in code. Let me load a few different timeseries into a dataframe. Most likely the easiest thing to do would just be to either embed numpy into your engine, or create simple wrappers to load data to it.
Supporting custom functions is definitely something we want to do.
double* values = (double*)malloc(sizeof(double)*num_of_values);
(If you are unfamiliar with numpy, it's just a python wrapper around raw blocks of memory.) double* data = shmget(key, sizeof(double)*num_data_pts, whatever_flag);
But directly embedding a Python interpreter is pretty easy in C, so I imagine it should also be fairly easy in Go as well.(Of course, given that my code snippets are C, you can probably deduce I've never written any Go, so take my comments on how easy it is with a grain of salt.)
Folks that wanted to could then write Python/Numpy code for these and use Numba+LLVM to compile them.
It could be quite performant and avoid having to marshal data or do IPC, could possibly avoid even copying data in some cases (Numpy/Numba have pretty robust support for structures coming over an FFI, not sure about FFI in Go.)
I think some more aggregate functions would be useful:
- Count (not listed on the Functions page, but used in the query examples?) - Sum - Standard Deviation - First - Last
It's definitely built with Go.
https://github.com/lsh123/stats-rrdb
An important part for me was the desire to completely separate data and UX (so no graphite). Added bonus is ability to control resources (e.g. memory/disk usage).
We run it in production for quite some time processing hundreds of data updates per second and tens of queries per minute.
http://sandbox.influxdb.org:9062/#/?username=ankit&password=...
Other than that, I look forward to evaluating this .. maybe its a solution for a problem I have recently where I'm collecting massive log files of operation systems, and need to navigate/parse/analyze .. so I guess I import the logs into InfluxDB, and put a d3.js frontend on it ..
Now not discouraging InfluxDB or anything, as a systems programming fan it's great to see more things like this coming, and as a Gopher too.
Good point, Comrade. I will propose to GOSPLAN that we rationalise the development of all new technologies, to avoid such accidental evolutionary convergence in future.
I know it's early days but I didn't see any information about cluster management - how does one setup an Influx cluster, can it be resized, what kind of hardware does it prefer?
The goal with the cluster stuff is that it should be possible to add nodes to the cluster, but the storage part of it isn't highly elastic. Meaning, you won't be adding and removing instances from it frequently. So adding nodes will require you to go into the admin interface, activate them, then wait up to half a day for rebalancing to be complete (but the cluster will be available for reads and writes during this time). However, we will be optimizing for the case of replacing a failed or soon to be shut down node.
If you're serious about giving it a try when we have the clustered version available, shoot me an email: paul@pauldix.net. Would definitely like to hear more about your use case.
My implementation would output a chart given parameters:
/chart.png?metric=whatever&time=12h&interval=10m
Are there any plans for easy output of graphs?
However, security is probably not something to bother with in sandbox since it's not HTTPS. We're looking at it now, and should have it back up in a bit.
What's the scalability model? It's not clear from the documentation.
[1]: https://github.com/influxdb/influxdb/blob/master/src/datasto...