Kerf: a columnar tick database for Linux, OS X, BSD, iOS, Android
github.com
github.com
The query language and concepts which are explained in the documentation are an exciting/novel (at least to me) contribution and "out in the open" for everybody to see/discuss/adapt/improve upon. Let's not dismiss that because we can't read the source.
Contrary to what you state, it is _not_ clear from the README that this is a commercial venture, just that the author provides an email address for any queries in relation to licensing, and actually includes _no_ copyright or restrictions on running the code in the provided documentation. One might even assume that since we've been provided a copy of the binary without any notices, we might be free to run said code without restrictions... but really, who knows? Hence the OP's questions.
Yes, just because it is hosted on github does not mean it is required to be open-source, but the documentation doesn't include any information about the licensing terms, so it is a valid question (that is as yet unanswered in public).
Also did you roll the SQL engine from scratch or does it built on some existing parser/runtime? Would love to learn more.
SQL engine: all from scratch
Are you planning on adding that and which syntax are you planning to use? Have you looked at postgres "OVER PARTITION"? While it seems powerful it's also fairly unintuitive IMHO. I was experimenting with adding a GROUP BY clause that allows each input row to appear in more than one group in the result set. Something like:
SELECT time, mean(value) FROM mymetric GROUP OVER TIMEWINDOW(time, 60);To make my question more precise; I was trying to ask specifically about a "moving window aggregation" (e.g. a moving average over a timeseries). This is more like asking the question "Please give me every minute an aggregate based on all values in the last N minutes". To do that you need each input row to end up in more than one bucket (or have a special type of aggregation function like postgres does).
For example, if you were doing a moving aggregation with a 1-minute interval ("bucket size") and a 5 minute window ("lookback"), you would need to place each row into 5 buckets: The bucket into which it belongs based on it's timestamp and the 4 previous buckets. And a vanilla SQL GROUP BY can't do that.
Hope that makes sense.
By method of exclusion, it cannot be musl because musl unifies everything into a single /lib/libc.so.
It most likely is not uClibc because uClibc is typically symlinked to /lib/libc.so.0 from /lib/libuClibc.
Considering that the soname of /lib/libc.so.6 has been reserved by glibc since the 2.x days (to differentiate it from the Linux libc days of libc4 and libc5), it is most likely indeed glibc in use.
It's obvious this program isn't making use of NSS or anything like that, so it's doable to use glibc for static linking, if frustrating.
ldd kerf-unpacked
linux-vdso.so.1 => (0x00007fff5e7f1000)
libpthread.so.0 => /lib/x86_64-linux-gnu/libpthread.so.0 (0x00007ff1833aa000)
libm.so.6 => /lib/x86_64-linux-gnu/libm.so.6 (0x00007ff1830a4000)
libreadline.so.6 => /lib/x86_64-linux-gnu/libreadline.so.6 (0x00007ff182e5d000)
libc.so.6 => /lib/x86_64-linux-gnu/libc.so.6 (0x00007ff182a99000)
/lib64/ld-linux-x86-64.so.2 (0x00007ff1835d1000)
libtinfo.so.5 => /lib/x86_64-linux-gnu/libtinfo.so.5 (0x00007ff182870000)
So using NSS shouldn't be an issue; and the binary actually does use it open("/lib/x86_64-linux-gnu/libnss_nis.so.2", O_RDONLY|O_CLOEXEC) = 3 libreadline.so.6 => /lib/x86_64-linux-gnu/libreadline.so.6 (0x00007ff182e5d000)
libreadline is GPL'ed, not LGPLed, Kerf is therefore in violation.https://cnswww.cns.cwru.edu/php/chet/readline/rltop.html
Edit: The GPL doesn't allow any linking into non-GPL compatible software, dynamic or static. Readline has this license for exactly this reason, and they have forced software into the public with it before (ncftp comes to mind)
http://www.gnu.org/licenses/gpl-faq.en.html#GPLStaticVsDynam...
Edit2: https://en.wikipedia.org/wiki/GNU_Readline#Implications_of_G...
EDIT: Looks like I stand corrected and this is NOT the case even if you are just dynamically linking a GPLed .so; check parent's links/ask a proper lawyer/please don't sue me ;)
I don't think anyone cares enough to sue you, however you are clearly in violation.
I'd like to look into it further, but I can't find any information about licensing. Given that there is no source code in the repo, it appears that this isn't an open source project.
"Contact Kevin (e.g., licensing, feature/documentation requests): k.concerns@gmail.com"
So if you're interested, I suggest you send Kevin an email.
Personally I think it would've been interesting if it was Free Software -- as this doesn't come with any license, it's essentially less useful than the 32bit kdb[1] -- and I can't imagine it's as feature complete, or has a similar level of documentation, support or real-world testing.
If the intention is to make the binaries deployable/usable, a license file/note in the Readme would be the absolute minimum requirement IMNHO.
Also find it strange to host just binary artefacts on github -- I can't see how that's useful for the author or the users. I suppose one could submit pull-requests against the python api, but then it would make more sense to split the repo in two -- one for the source-available (unknown license) python code -- and host the binary artefacts somewhere.
"Tick databases are real-time database servers for capturing, managing, and processing market data."
They go into the requirements after that.
More generally, tick databases are databases specialized for storing and querying high frequency time series data.