Yandex open sourced it's BI tool DataLens
github.com
github.com
I work on an open source code-based BI tool called Evidence, which might be of interest to you.
It's effectively a static site generator aimed at building automated reports and analysis.
https://github.com/evidence-dev/evidence
Previous discussions on HN:
https://news.ycombinator.com/item?id=28304781 - 91 comments
https://news.ycombinator.com/item?id=35645464 - 97 comments
I want to be able to publish this notebook when I’m done and then be able to hand that around, the same concept that you’ve built Evidence around. I think that’s a very good idea, so thanks.
The main thing keeping me from switching is that Metabase’s query builder and visualizations are too good for 95% of my work. It’s hard to picture going back to writing _all_ the SQL.
I hear that. We're making a lot of progress on reducing the amount of sql you need to write, keeping it DRY etc. Making the dev experience really buttery and high leverage is definitely a priority.
Here are a few of the things we're working on in that regard:
1. Making our components issue their own queries so that you don't need to write full sql expressions, just fragments when you're defining the chart you want.
2. Improving intellisense -- right now you get "slash commands" and snippets in our vscode extension to invoke components (which are really sweet), but we're aiming to get to a full intellisense type of experience.
3. Supporting external metrics layers where it makes sense. We've got some users using Cube, we're interested in Malloy, and the dbt semantic layer, those types of things.
One of our team members is very keen on building something he calls "Evidence studio", sort of a wsywig you could invoke during development for generating basic queries, setting up charts etc. that syncs the components into their text representation. That'd be further off though :)
For anyone who is interested in our cloud service, it's an easy way to put your project online, keep it up to date with your data, and place it behind access control.
For many organizations, hosting Evidence in their own infrastructure is easy enough, but if you don't want to manage that, we are happy to manage it for you.
It is not free (that's how we pay the salaries), but pricing is available here:
Superset looked good, but operating superset quickly runs into the same Python issues all Python software suffers from.
Sometimes it would just break for no apparent reason. Configuring it was a nightmare of magic Python code and unclear settings. Trying to use plugins was equally painful: due to the poor boundary separating the applications dependencies from the plugins dependencies, adding a db connector could just bork the whole application.
https://github.com/apache/superset/issues/13345
Almost instantly run into this issue setting up a test instance of Superset. And the issue has been around for years.
I was thinking of something closer to Power BI, e. g. something like AppSmith but more BI-oriented (AppSmith is a generic tool to build your own applications, kind of like Visual Basic or Delphi).
Also CubeJS if you want to a bit more flexibility.
My preferred approach is implemented in Zillion, which I use for BI at my company: https://github.com/totalhack/zillion
it's => its
>We use the .us-data folder to store PostgreSQL data permanently. You can delete this folder if you want, it will be recreated with the demo data after restarting the datalens-us container
Why "us"?
I really like the concept of Domo. They have ETL, modeling, a warehouse and BI in one app ("data-stack-in-a-box"). I've interviewed 20 of their customers and the general sentiment was pretty bad. There's a long sales process, a longer process to get it set it up, and they've built all the modeling and connectors themselves (vendor lock in, none are best-in-class).
Definite (https://www.definite.app/) is a data-stack-in-a-box. We have a built-in modeling layer for core metrics and an AI assistant to answer any one-off questions.
A few ways we're different:
Built on open source - We run the data stack for you and give you a single app to manage and analyze your data, but it's all built on open source standards. So if you decide at any point you want to run it all yourself, the code is yours to lift and shift to your own infrastructure.
Battle tested connectors - We're using Meltano / Singer (open source library from Stitch) for our connectors, so they've been used heavily in production for years.
Self-serve that actually works - A lot of tools promise self-serve, but AI is making this real. We've invested heavily in making it possible to ask questions and get accurate answers. The AI queries a modeled view of your data that can answer questions that depend on well defined metrics (e.g. ARR, DAU, etc.).
Can you share any information about how it knows where to join?
1. When you bring your own data warehouse, we parse the query history, convert it so an AST and learn JOIN's from there
2. When you're using our managed data warehouse and ETL, we already know most of the JOIN's (e.g. we know how to join the Hubspot data we ingress to Stripe)
3. For anything not covered by #1 or #2, we have a modeling layer where you can specify how to join tables.
https://www.intellinews.com/russian-tech-titan-yandex-ceo-vo...
It's a shame that geopolitics means most of it will have to be reinvented by someone else before it'll see any use.
I can't see that argument working in a court of law...
Don't US(or any other country for that matter) laws discriminate based on nationality all the time?
Yandex is aware of how the geopolitical situation is hurting them and are therefore building a new company called Double.Cloud, based in Europe, to work around the negative public opinion on Yandex, and thereby continue being able to sell Clickhouse cloud services.
I can buy or manually provision anything, no technical hurdles or policies from that side. My absolute focus is the raw UX for business people.
Suggestions?
We designed it specifically to provide an excellent UX to business users while reducing BI burden on the data team. We find that most business users often just need to search, filter, and sort instead of looking at charts to make operational decisions.
UX-wise, what sets us apart are:
- <1s full-text search (even on billions of rows of data), feels like Cmd+F in Google sheets, but faster
- Performance: we stream billions of rows into the web browser, seamless scrolling (no paging of 50 records at a tieme)
- Rich cells make tables easier to scan/read (enum strings => colored tags, numbers => color-coded based on value => checkboxes, timestamptz => clear date time pills)
If that fits what you need, happy to give you a demo.
arthur(at)dataland.io
Otherwise, I think the simplest BI (if charts are impt) could be something like evidence.dev or Metabase.
But I also think it's going to require some curation on your part. Can you reasonably expect business users to navigate the entire schema/table tree across these three sources? That's where I think the bottleneck often lies -- if your BI tool allows engineering to just expose a subset of curated core tables.
Maybe that'll change one day with AI, and when it does that will be bought by every big company in the world (-:
Metabase is the only tool I’ve used where I’ve managed to get non-technical users to actually engage and use to query building tools to answer their own questions.
Follows a conversational "ChatGPT-like" approach since already 2016.
Info: I'm one of the founders.
> has the best UX for nontechnical people to assemble some data
If they can use Excel / pivot tables, they can use Definite. They can also just ask in natural language and we generate the report for them.
> Salesforce, some mariadb/postgres and (optionally) hubspot as data sources
We have pipelines for all of these and can spin up a managed data warehouse to store all the data if you don't already have one.
Drop me a note at mike@definite.app if you're interested
ps - I spotted the pedant that is technically correct, and I claim my 5 McFun bucks!