Apache Superset
superset.apache.org
superset.apache.org
Superset allowed us to replace Tableau and not looking back
Took me a while figure out how to embed it into my app using Superset Embedded SDK.
Superset Embedded SDK - "Embedded SDK allows you to embed dashboards from Superset into your own app, using your app's authentication. Embedding is done by inserting an iframe, containing a Superset page, into the host application."
https://github.com/apache/superset/tree/master/superset-embe...
Superset is based on very high quality and well maintained chart library eChart
https://echarts.apache.org/examples/en/#chart-type-linesG
Community Roadmap
https://github.com/apache/superset/projects?query=is%3Aopen
Huge respect to Preset.io and its team for contributing to the project and keep it in a great shape
Superset source code is very easy to read and understand, and as a result it's possible to implement some advanced caching techniques reduce the load on charts.
No BI is perfect.
Watching Superset for years gives me confidence the project will work as supposed down the road, and eventually some of its packages can be reusable for all kind of visualizations and data hacking.
Our main approach to visualisation is to start with eChart and simple Reactjs wrapping and spin off Superset on subdomain for power users, and later see which one works better. Same look gives a very pleasant experience.
https://www.tableau.com/products/server
If you have money, dedicated team of data analytics who are already familiar with Tableau - no need to torture them with other tools.
Superset is lightweight and open source, but only has 5% of the features. So it really depends what you need!
Their interactive "embedded-mode" avoids iframes too... but it's built with web components, so you wind up in shadow-DOM hell if you want to do anything dynamic on the view's contents.
Previous HN discussion: https://news.ycombinator.com/item?id=35645464 (97 comments)
Reminds me Obsidian DataView but with charts https://github.com/blacksmithgu/obsidian-dataview
This whole ideas to have data, visualisations and knowledge base in one private offline place is very appealing
The Markdown <-> Markup typing experience is just so good compared to e.g. Slack, Reddit and other markdown-esque tools
Here's an example project with some filter components and custom styling: https://ecommerce.evidence.app/
This is still a static app - the data warehouse was only hit during the app's build process
All in all, not very innovative, but highly needed open source version of a traditional BI tool. Definitely something to follow and to use in temporary, not too demanding use cases. And hopefully a future replacement of Tableau or Power BI.
I tried Superset a few years back, and maybe it's changed since then, but intuitive is about the last thing I'd use to describe it. Things which I could figure out in a few minutes on any other BI tool literally took me hours of searching. It didn't help that they decided to rename core concepts at some point so half the online documentation made no sense anymore. Others at those companies who tried it at the time said similar things.
Demo is nice.
Instead I installed Metabase in 5 minutes tops: spin ec2 instance, whether and java -jar . I've never looked back.
The only thing that turns me off I'd that it's implemented in an obscure language. At one time I wanted to add some custom postprocessing to an api (given an sql query, get some python/pandas postproc command from a sql comment and execute it in the returned table), but the used language is just not for me (some lisp dialect)
I like Grafana too, but there's basically no isolation between your query and the SQL database at least in the Altinity Grafana plugin for ClickHouse which is the main one I use.
And the documentation is sparse at best.
1. https://www.apache.org/foundation/how-it-works/#incubator
What it was though, was riddled with dozens of Python runtime errors and innumerable glitches.
Metabase is where it’s at.
1. Built-in data warehouse - We spin up a duckdb database for you to load data to
2. 500+ connectors - You don't need to buy a separate ETL and you can pull in all your data (e.g. Postgres, Stripe, HubSpot, Zendesk, etc.) automatically
3. Semantic layer - Define dimensions, measures, and joins in one place. We have pre-built models for all the sources we support (e.g. the Stripe model already has measures for MRR, churn, etc.)
4. Simple BI - Build a table with the data you want and generate visuals off that table
I'm mike@definite.app if you have any questions.
But they may be an overkill if your primary use case is to infrequently build semi-interactive reports for non-technical end-users and your use cases are are mostly covered by standard graphs & tables. Esp. so if you are familiar with SQL and have access to the underlying data source. Two nifty utilities I have found to be very useful for latter kind of use cases are SQLPage and Evidence.
They make it very convenient to whip out some SQL and convert that to a neat professional looking web ui that can be forwarded to an end user. In case of Evidence it is a statically generated site, and in case of SQLPage it is a web app that connects to a live database.
SQLPage: https://sql.ophir.dev/
Evidence: https://evidence.dev
I've been running it in production since 2017, at two jobs, the current one a big corporation.
Best general-purpose, database-backed dashboarding system out there. I would never pay for Tableau or PowerBI.
Same for Airflow.
https://phabricator.wikimedia.org/T169452
Back then, I used this to generate some custom statistics
Open source Business intelligence platform made with Python - https://news.ycombinator.com/item?id=29368664 - Nov 2021 (49 comments)
Apache Superset 1.1 - https://news.ycombinator.com/item?id=27439939 - June 2021 (28 comments)
The Apache Software Foundation Announces Apache Superset as a Top-Level Project - https://news.ycombinator.com/item?id=25905277 - Jan 2021 (1 comment)
Apache Superset is an enterprise-ready business intelligence web application - https://news.ycombinator.com/item?id=21133931 - Oct 2019 (7 comments)
Is it worth it for BI on small datasets?
The reason we chose Metabase was that it had table joins, while Superset doesn't (unless it has added them since I used it). It also looks a bit sleeker. But I strongly prefer Superset; I found that with Metabase I had to turn a lot of things off to make it usable (Let me see "the_table" not "The Table"!), I was constantly annoyed at the opacity around models vs "questions", etc. and every time I wanted to change a question Metabase insisted on creating a new one instead. The real issue here was when we wanted to swap out the data source for a lot of questions but there was no clean way to do so without MB just creating new questions.
Also, Metabase doesn't have serialization unless you pay them AND you self-host, (if I'm self hosting then what exactly am I paying for?) and that's pretty annoying. https://www.metabase.com/docs/latest/installation-and-operat....
But it does let you join tables. Sometimes that's enough to make MB worth dealing with.
Yeah, somewhere along the line Metabase decided to get opinionated on "self-serve". I imagine it works well for some teams and companies, but for the tech-oriented, it's annoying.
I prefer my BI tools to be platforms that make for easy charting and cross-filters, while I build and control the models behind the scenes with a tool like dbt.
I've found the weird "make it easy" mindset a bit annoying with Metabase too. The whole questions, nice table names...
I'll give Superset a try in my next project I think.
I recently ran a little shootout between Superset, Metabase, and Lightdash — all open source with hosted options. All have nontrivial weaknesses but I ended up picking Lightdash. Superset is the best of them at data visualization but I honestly found it almost useless for self-serve BI by business users if you have existing star schema. This issue on how to do joins in Superset (with stalebot making a mess XD) is everything difficult about Superset for BI in a nutshell. https://github.com/apache/superset/issues/8645
Metabase is pretty great and it's definitely the right choice for a startup looking to get low cost BI set up. It still has a very table centric view, but feels built for _BI_ rather than visualization alone.
Lightdash has significant warts (YAML, pivoting being done in the frontend, no symmetric aggregates) but the Looker inspiration is obvious and it makes it easy to present _groups of tables_ to business users ready to rock. I liked Looker before Google acquired it. My business users are comfortable with star and snowflake schemas (not that they know those words) and it was easy to drop Lightdash on top of our existing data warehouse.
(one of the maintainers of Lightdash) You touched on some of our most interesting problems here! Would be especially interested to hear about what you liked / didn't like about symmetric aggregates in Looker and how you find dev with YAML. If you have an idea of how you'd like these to look in Lightdash, the team would be really open to making that a reality.
For pivoting in the backend, this is coming! Issue here: https://github.com/lightdash/lightdash/issues/2907
Another option could be to use LLM to summarize, tag and group queries for better discoverability.
One thing that helps is hooking metabase up to its own database and building queries on your queries, e.g.:
select *
from report_card
where dataset_query ilike '%' || {{query}} || '%'
(You can also join in metadata like the author, when it was last ran, etc.)We also try really hard to keep the Collection directory structure clean and consistent. But it's still really hard.
Any recommendations for a good piece of software for the single user case? Or a more convenient way to run the heavyweight tools?
This is a single user application, unless you make it part of your built application.
K8s installation instructions: https://superset.apache.org/docs/installation/running-on-kub...
RBAC configuration: https://superset.apache.org/docs/security/#rest-api-for-user...
But if your vis are with the scope of native Tableau capabilities, then Tableau it's more friendly and gets less in the way of you and your work.
I see these products as tools for data visualization and reporting i.e. presenting prepared datasets to users in a visually appealing way. They aren't as well suited for serious analytics.
I can't comment on Superset or Tableau but I am familiar with Power BI (it has been rolled out across my org), the type of statistics you can do with it are fairly rudimentary. If you need to do any thing beyond summarizing (counts, averages, min, max etc). It is not particularly easy.
For data analysis I use SAS or R. This software allows you do things like multivariate regression, timeseries forecasting, PCA, Cluster analysis etc. There is also plotting capability.
Both these products are kind of old school, I've been using them since early 2000's, the "new school" seems to be Python. Pretty much all the recent data science people in my organization use Python. Particularly Pandas and libraries like Seaborn (https://seaborn.pydata.org/).
The "power" users of Power BI in my organization tend to be finance/HR people for use cases like drill down into cost figures or Interactively presenting KPI's and other headline figures to management things like that.
I am in need of a "dashboarding" feature in our SaaS, but it seems there's a gap between PowerBI/Tableau/Metabase/Superset and various charting libraries. The former are too much "turn key" and the latter require a ton of work to setup all the chart-building UI and features...
It’s commercial software though.
It's part of our Open Source Data Platform and it's one of the few open source BI tools out there and there are not a lot of alternatives in this space. We generally like it.
Many, or most, users for a BI tool will be operations, product managers, and business management who simply will not find the interface to be intuitive, responsive, or well designed. At least that's my experience.
Wes McKinney used to have an excellent 5 minute introduction to pandas in this genre.
AFAICT needs a db (MySQL/Postgres) and a cache (Redis/Memcached) and one (or more?) web workers.
Then optionally also Celery workers (for "async queries" i.e. slow running)... not sure how optional that is though.
If you want to style the whole application, you can fork the repo and go bananas. If you're looking for theming, there's more to be done yet on that front, and I wrote an article on that too: https://preset.io/blog/theming-superset-progress-update/
[1] https://blog.adnansiddiqi.me/create-your-first-sales-dashboa...
Superset Embedded SDK - "Embedded SDK allows you to embed dashboards from Superset into your own app, using your app's authentication. Embedding is done by inserting an iframe, containing a Superset page, into the host application."
In what way, any details?
Not been tracking that or using OpenOffice for a while.
As such, closer to an open source replacement for PowerBI.
superset is the flip side of grafana; not good for up-to-second updates, but good for complex queries. Also, non-time series stuff. Ex: Which customer groups bought which products for all time? <— that type of BI stuff.
Grafana I sense is culturally focused on observability visualization (aka needs of full stack devs). Culture is very hard to change!
To make it more concrete -- coworkers tell me Grafana doesn't work so well with Apache Druid, while Superset supports it quite well.
Apparently installation will not work with Python 3.12, dur to deprecation of distutils.
Does anyone have any method to install this?
https://superset.apache.org/docs/installation/installing-sup...
we have multiple tenants + developer instances of our warehouse. to reuse the same dashboard in this setup we need to create at least 3 virtual datasets, plus wrangle a bunch of boiler plate jinja.
I think Athena can only query data on S3?
https://preset.io/blog/accessing-apis-with-superset/
"Shillelagh (ʃɪˈleɪlɪ) is a Python library and CLI that allows you to query many resources (APIs, files, in memory objects) using SQL. It's both user and developer friendly, making it trivial to access resources and easy to add support for new ones"
I wrote about Superset's semantic layer here: https://preset.io/blog/understanding-superset-semantic-layer...
One popular option is to use dbt or Cube for the semantic layer and pair with Superset: https://preset.io/blog/announcing-presets-ui-integration-wit... and https://preset.io/blog/open-source-looker-cube-superset/
I built my own semantic layer instead. I use this in production in my company but obviously use at your own risk as it's a one-man show.
There is this:
> A final SQL query against the combined data from the DataSource Layer
> The Combined Layer is just another SQL database (in-memory SQLite by default) that is used to tie the datasource data together and apply a few additional features such as rollups, row filters, row limits, sorting, pivots, and technical computations.
But it leaves me with questions - how/when does this get populated? What other options are there besides in-memory SQLite? (I presume that's just a convenience for development and would use something else in production?)
Or is it just what Superset calls a 'metastore' i.e. data about the data, and the queries are run against the data source layer?
> Superset lets you join tables within the same database. If you want to do cross-DB joins, we have a new (beta) in-memory meta-DB that lets you do this
...is it this?
It first runs one or more queries against your DataSources in a drill-across query fashion. You can think of DataSources as one or more completely separate databases. You could have one mysql, one postgresql, one duckdb etc all in the same Warehouse (not saying this is common in production, just an example). Within those DataSource queries it's also joining all needed tables together for you, i.e. joining multiple tables in each database to meet your required grain.
It then takes the results of all those queries and combines that data in another layer which is currently an in-memory sqlite database. The purpose of that layer is joining the data for presentation as well as applying some additional features like rollups, technicals, formula fields, etc.
I'm not familiar with what superset does under the hood or exposes as an API so I don't know how to compare it, if there is some similar backend piece. But I suspect no part of superset is quite the same as this, based on what its front end can do.
Happy to answer any other questions you have.
So now that I picked Metabase, Superset is topping HN for no apparent reason. Why?
Amen brother.
Can you name a few examples?
OpenOffice is probably the most famous (it still has the name, but it is dead, LibreOffice is the real "active" fork).
And the things in the "Attic" are officially dead - https://projects.apache.org/committee.html?attic and many more projects should be there.
Lots of other projects just die silently and/or you are unsure of the status.
Here you at least have a chance to revive them if you like as there is always an overarching organisation.
More things should move into the Attic, like OpenOffice.
https://projects.apache.org/projects.html?name
These are all projects that once were (more) relevant, however seem to have become rather niche (Gradle, Jetbrains/VSCode, GoogleDocs/Libreoffice e.g. for the first three are the dominant competitors).
Most of these projects (like the massive commons listings) are either used by some Java library somewhere (meaning their success/relevance is tied to the usage of Java), or are obscure enough that they are no longer used widely and so suffer from lack of interest.
There are gems in this list, to be sure, but if you just run into half-maintained projects all the time you're not likely to associate good things with the Apache name?
Hmm. I suppose all open source looks that way if it doesn't get regular funding/attention.
Apache does house a lot of abandonware. They had some relevance as recently as 6-7 years ago but they've been largely replaced by nginx I think. That being said, I view them like the local soup-kitchen - important to have and maintain, but not where I want to go for a 5-star meal.
These projects don't give the apache foundation an appearance of importance or relevance, rather they make it look rather rundown.
Obivously the main "alternative" is for the original company to simply shut down the product/service, which can do irreperable harm to a company when they have high-profile customers who are utterly dependent on a service.
Another alternative is to use an open-source foundation that's directly managed by the original company, which is what Microsoft did with its DotNet Foundation ( https://dotnetfoundation.org/ ) - and while Microsoft's legal team ensures the foundation is "legally" independent, in practice we know all the significant shots are being called from within Microsoft-proper; but it does give us some modest reassurances that .NET won't suddenly return to being closed-source overnight.
Another alternative is to not open-source it and to instead sell it off to another company that can maintain it while still being profitable - this is what Adobe did with Flash: they sold it all off to Samsung because their Harman division wanted to continue using Flash for embedded/automotive UX work. This approach can work, but doesn't benefit the wider ecosystem the way that open-sourcing does - and something something shareholder value and return-on-investment by selling rather than writing-it-off...
What companies won't do is let any of their engs that are passionate about a project split-off from the company to run and maintain it, le sigh.
I think the HN algo is pretty easily manipulated. I worked at a startup that had an effective process to get things to the front page
That sounds (potentially) sleazy. If you think it's a technique that HN could potentially defend against, I encourage you to explain it to hn@ycombinator.com.
Pretty sure it's as simple as posting in your general slack channel "@here we posted a new article to HN, go upvote and write a comment"
I don't think it's semi-abandoned. I had a brief interaction with the project in my previous job, and I found the community and the company to be reasonably engaged and responsive.
Business users loved the self-serve query builder, and it wasn't uncommon to walk around the office and see Metabase up on someones screen. My CEO absolutely loved it, and used it daily including to put together data for board decks.
None of my users cared about visualizations, and lived in tabular data. This included finance, marketing, merchandising, operations, and executives (CEO/COO/CFO). The only people that lamented the limited visualization were analysts. Power users did all their day-to-day work in Excel or other tools anyway, such as managing marketing spend or inventory allocations.
Metabase was great for dashboards and self-service (ad-hoc). 10/10 would deploy again.