Show HN: Open-Source Business Intelligence for BigQuery – Looker Alternative
mprove.io
mprove.io
But more importantly, the challenge for any such tool is to go beyond use by 2-3 people. At 2-3 people anything will work. Where BI tools (open source and close source) struggle is scale: having all the right features for, essentially, a group of users who actually don't know how to work with data (did I just say that aloud?). Chartio caps at 20 people. RJ capped at 50-100 (and later became Stitch for that reason). We haven't seen where Metabase caps, but I bet it is in a similar range. Very few BI products have actually surpassed 100 users at target installations. And beyond 1,000 is a real challenge that only few, and even then with a lot of assistance, can support: Tableau, Looker, Microstrategy, maybe Birst, maybe Domo.
Also, a combination of BI with LookML is a complicated product. During my days at Looker, we were handling 50+ bugs / week, and filing 1,000+ tickets. Every day we were filing over a 100 new features.
So with all that, the question is, is it really worth the struggle? What's the end vision for supporting this? Why should someone who implements BI for a living bet on this product?
Also, if you think this is a rip off of LookML, you should take a look at what GitLab is doing with Meltano. They completely jacked LookML.
Is it something technically related or more business/feature related?
Tech-wise, data stacks are complicated. Hundreds of pre-existing tools, with many owners and diverse interests. BI is where they all meet. With some organizations its easy to build a central data warehouse, but in a large corporate environment that's a dream. Imagine how many data warehouses an organization like Microsoft might have? OK, MS is too large - how about Expedia or Zillow? Thousands? So you connect a tool like Looker--that was designed for 2-10 connections--to a thousand? How do you even administer all the connections? How do you govern access? Access permissions at Looker (and other BI tools) are great, but it would take you 10 years to set them up for 1,000 DBs. And what about tool access? Some want SSH tunneling, others want complete on-prem. Some need Google authentication, while others want one through a pre-existing corporate login. Most large organizations typically end up building--or hiring someone to build--their own custom solutions on top of these tools. Such custom solutions are bigger projects than the implementation of a BI tool to begin with. Oh, and most also fail. The success of a BI in an enterprise environment is partly good sales/marketing/support, but a big part is due to the product achieving some form of maturity with all this side tech--the stuff totally not core to the analytics itself.
Business-wise, try getting people with diverse backgrounds agree on common terminology. Let's take marketing for instance. Should be easy, no? But hey, marketing is actually like 10-20 different kinds of people - some know SQL, others can't put 2+2 together and arrive at anything other than the word "magic". Some think in terms of stories--others think in terms of conversion funnels. OK, so you've put together some data dictionary, did some training, segmented users into 1) technical ones (SQL/LookML/BackML...), 2) business (explore), 3) consumers (dashboards/pretty charts). But now it turns out that much of the data they rely on is generated by a different team--say, operations or sales. Again, you've got 10-20 different kinds of personalities there. Somehow all these people have to agree. How do you make them all agree? Short answer: you cannot. No one can. They don't even like speaking to one another - and now you are going to come in and make them agree on what kind of KPI is going to determine their success => bonus. Hell no!
There are ways to make progress on both fronts. And, no doubt, an open source project has some chance. But it is not easy. And no one has really done it in 30 years. Many have tried.
Full disclosure: I love BigQuery (Google Cloud partner). Love Looker. Rely on Open Source constantly. Just trying to demonstrate a realistic view of how hard this problem is.
Again, if you don't mind sharing, what would be a dream solution for you, something that could solve most of the problems you mentioned?
I understand that some of them are cultural problems (like people no talking to each other), but maybe somehow a technology could solve this?
If you could think of a new approach, what would that be?
For analysts who work with SQL queries, any wrapper on a language seems complicated. But in the case of LookML and BlockML, in a couple of days you already begin to understand how this works, and after a week you can not tear yourself away from using the product.
Looker and Mprove are in the middle between fully custom DIY solutions and simple query generators tools.
I met the online Looker demo a few years ago when I was looking for a business intelligence tool for another project.
Looker has a closed source code, so I did not see what algorithm is used to build queries.
I kept those LookML features that I understood and liked. However, in some places LookML is confusing (for example - references and aliases). I made them differently.
Later Looker quit using YAML.
> Very few BI products have actually surpassed 100 users at target installations. And beyond 1,000 is a real challenge that only few, and even then with a lot of assistance, can support: Tableau, Looker, Microstrategy, maybe Birst, maybe Domo.
Scaling that size is not a top priority right now.
> Also, a combination of BI with LookML is a complicated product.
You should have told me this a couple of years ago. It seemed pretty simple.
> So with all that, the question is, is it really worth the struggle?
Yes, If people will use it.
> What's the end vision for supporting this?
In future - maybe "skinny" or "thin" option mentioned here - https://medium.com/open-consensus/2-open-core-definition-exa...
>Why should someone who implements BI for a living bet on this product?
If you can not afford Looker (like me) and want to use similar product.
Congratulations on the launch. I will be using mprove and will give you feedback along the way.
EDIT: I would love to get in touch with you and learn more about what your plans are. Possibly collaborate if you'd be open to that. I'm dpaola2@gmail.com -- I couldn't find your contact info anywhere!
At one point Looker was essentially a web application with a query service that could be scaled by adding more servers behind a load balancer. Each end user is essentially working in a shared nothing environment and the query engine is driven by metadata stored in a git repository. Looker itself did not manage cache or analytical processes. All of the real effort to scale was in the database backend.
In other words: Looker ought to embrace the idea of LookML being the data modelling language. Encourage other implementations of it. Treat them as customer sources instead of customer sinks. That's how they can take over the world :)
Also, it's definitely a challenge to support 100's and 1000's of users digesting data, especially in the democratized fashion that we're typically used in. It takes good data governance, support, and admin tools. I gotta chime in and say for the record though that Chartio well supports many such customers.
It may be weird for you to chat as you worked at Looker, but I'd love to hear anytime on why you see Chartio capping out at a lower # than the others listed. You can reach me at dave-at-chartio.com!
Quick point however - why do you need a new database ? You can use a table inside bigquery itself. It seriously reduces the dependencies required.
But you can't use BigQuery for OLTP.
https://github.com/mprove-io/mprove/blob/master/deploy/docke...
Can you not use biquery database itself. Create tables for your internal use instead of MySQL ?
I would take higher latency, but avoid pulling in a whole database infrastructure.
Plus a huge number of us use postgresql..so that becomes another set of a mess. I would strongly urge you to do this on the same bigquery database that you would connect to anyways.
Here's a recently built Graphite connector as well:
/disclaimer: work in google cloud
CREATE TEMPORARY FUNCTION mprove_array_sum(ar ARRAY<STRING>) AS
((SELECT SUM(CAST(REGEXP_EXTRACT(val, '\\|\\|(\\-?\\d+(?:.\\d+)?)$') AS FLOAT64)) FROM UNNEST(ar) as val));
... SELECT COALESCE(mprove_array_sum(ARRAY_AGG(DISTINCT CONCAT(CONCAT(CAST(a.id AS STRING), '||'), CAST(a.population AS STRING)))), 0) as a_cohort_size
Recently, Looker began to do it differently, most likely to improve bigquery performance: COALESCE(ROUND(COALESCE(CAST( ( SUM(DISTINCT (CAST(ROUND(COALESCE(lesson_5_cohorts.population ,0)*(1/1000*1.0), 9) AS NUMERIC) + (cast(cast(concat('0x', substr(to_hex(md5(CAST(lesson_5_cohorts.id AS STRING))), 1, 15)) as int64) as numeric) * 4294967296 + cast(cast(concat('0x', substr(to_hex(md5(CAST(lesson_5_cohorts.id AS STRING))), 16, 8)) as int64) as numeric)) * 0.000000001 )) - SUM(DISTINCT (cast(cast(concat('0x', substr(to_hex(md5(CAST(lesson_5_cohorts.id AS STRING))), 1, 15)) as int64) as numeric) * 4294967296 + cast(cast(concat('0x', substr(to_hex(md5(CAST(lesson_5_cohorts.id AS STRING))), 16, 8)) as int64) as numeric)) * 0.000000001) ) / (1/1000*1.0) AS FLOAT64), 0), 6), 0) AS lesson_5_cohorts_m_sum_distinctI can't seem to find any on the site...
I assume there was a time when this was meant to be a non-open source project.