Data Brewery (open-source data processing + OLAP in python)
databrewery.org
databrewery.org
As for Cubes: goal is to create light-weight framework with pluggable backends. Currently simple SQL backend and MongoDB backend are implemented.
Some public projects that are using Cubes for OLAP are:
Donations for sport and culture:
http://granty.transparency.sk/en/
Public procurements of Slovakia (still under development):
http://vestnik-test.democracyfarm.org/en/report/all?cut=date...
If you are asking about performance, my answer is: I do not know yet, haven't stressed it too much. I would very like to hear any feedback and/or recommendations. Current focus was on simplicity and easy of use, performance will come later.
For brewery, here are some blog notes:
Presentation where data brewery was used in a project:
I hope to prepare more information soon, with examples. I want data brewery to be more distributed with cusomisable nodes (like you would be able to use a distant server as a processing node or part of processing stream).
Goal of data brewery is to provide "way of working with data streams", focusing more on data analysis than on data transformation. However, it does not mean that you would not be able to use it for the further.
Anyway, I would appreciate any feedback, and gladly answer any questions. I am also looking for cooperation, if you are interested, drop me a line.
Stefan - @Stiivi on twitter, author of Data Brewery/Cubes
One thing I wonder is how easy it would be to integrate this with a MongoDB or Redis backend.
Other useful links I found on my quest:
- https://github.com/rsim/mondrian-olap (jruby olap queries on mondrian)
- http://www.slideshare.net/rsim/multidimensional-data-analysi...
- http://www.amazon.com/Pentaho%C2%AE-Solutions-Intelligence-W...
* What kind of limits and performance does this implementation have?
* Will data be fetched from the database for each query?
* Is it possible to have dimensions with millions of
values and expect reasonable query times?
* Looks like it supports advanced topologies and hierarchies.
How will dimensions with a high carnality affect performance?If you find out, I'd like to know too.
Before I answer your questions (I assume that you are referring to Cubes - OLAP framework), I think it would be good to note, that Cubes has pluggable backends. Currently simple denormalisation-based SQL backend and MongoDB backend are implemented. I want to have them more advanced.
* Will data be fetched from the database for each query?
- currently yes, however we did some experiments with plain HTTP caching of Cubes/Slicer server and it worked pretty nicely for our current needs
* Is it possible to have dimensions with millions of values and expect reasonable query times?
- not tested yet
* Looks like it supports advanced topologies and hierarchies. How will dimensions with a high carnality affect performance?
- right, it supports hierarchies, however same as above: not tested yet for performance
I am open to any commeents/suggestions regarding the framework(s).
Stefan Urbanek, @Stiivi on Twitter (author of Cubes)
@Stiivi - author of Brewery/Cubes