Do we really need a third Apache project for columnar data? (2017)
dbmsmusings.blogspot.com
dbmsmusings.blogspot.com
On nginx, this was going to require either recompiling the whole server to enable to some non-default module, or delegating the auth to a separate script/process that appeared like it was going to be painful to set up, or implementing the whole thing in lua. On Apache, it was a 10 line config file and I've never had to think about it again.
(And yes, I know the article is about the ASF and not httpd, but I just wanted to respond to this comment in particular.)
On a tangent, are there any other series these days like Andy Pavlo's Quarantine DB talks? I have a really hard time finding the covid equivalent of meetups. Most local meetups seem to have just stopped rather than go online.
Unfortunately I’m not aware of any other similar series. I’d love to find more content like this though.
Oh look, yet another Apache real time/batch/big data/stream processing/ingestion/workflow/whatever product.
Apache Druid
Apache Spark
Apache Storm
Apache Flink
Apache Beam
Apache Apex
Apache Airavata
Apache Samza
Apache TEZ
Apache Hama
It's basically a terrible joke at this point. There's no single Apache page helping you to decide which one you want, and they all seem to have such large overlap. Most of them seem to have bad documentation, and give the appearence of not really being maintained. This puts me off even trying to use them. If there's this much scope creep/NIH/reinventing the wheel happening across the board, I can't imagine how bad each product is individually.Apache Kafka seems to be the only exception...
Apache Pulsar Apache Pinot
They're not a corporation with unified goal, just an umbrella org.
https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/burstabl...
Is anyone aware of 2-dimensional storage hardware?
(It's so-called five dimensional because you have polarization angle and color in addition to physical position)
What a column-store database gets you is fewer I/Os on reads since all the data for a column is stored together. This is ideal for some analytic queries on warehouse databases with wide rows.
https://www.postgresql.org/docs/12/indexes-index-only-scans....
Said another way: in analytic workloads the hard part is not finding the data, it is reading the data.