Is BI dead? – On dismantling data's ship of Theseus
benn.substack.com
benn.substack.com
Firstly Tableau, QlikView or PowerBI are all pretty much doing the same thing whatever flavour you prefer.
We find that maybe 10% of users will actually use them to garner new actionable insights and 90% will moan there's too much data for them to process and no one appreciates the immense amount of data munging that goes on behind the scenes to make those pretty graphs.
Where do we go? Personally I see a lot of the insights being automated and turned into actions on the server side without anyone touching a BI tool. This can be achieved with rules based approach and perhaps some correlation, trend analysis over time series data. Where you do need to go beyond that then perhaps AR/VR will provide a novel and more valuable approach.
I reckon of the 90% that don't make any data driven decisions with the dashboards they're provided, there are three causes I run into frequently.
1. They can't describe what info they need
2. That info isn't available.
3. They don't make decisions based on data at all.
I feel I'm getting better at addressing 1 and 2 the more I do it, but I figure there's a fair amount from group 3 who demand the reports nonetheless...
As an aside, you mention Tableau, QlikView or PowerBI, have you tried Apache Superset at all?
It’s a very important function: provide data that can be cherry-picked to support a decision already made.
If the decision is too complex for that, hire a big consulting firm to write a report suggesting and justifying the action desired.
A decison staple surely going back millennia.
One way to solve is to automate the action from insight part and inject into users workflow. I am looking at this now with support of senior sales stakeholders.
Simple example:
Client X activity has been trending down vs it's peer group, Call Client X to understand what's going on.
Only works if you've got buy-in from management though.
Nobody is saying "I'd love another tool that I am expected to check every day", both embedded analytics and insights are how to drive adoption and unlock value from data.
It feels like a good idea, but I imagine it would be very fragile in black swan events.
Maybe a mix?
Provide the dashboards, but allow users to set up alerts/workflows, they understand the problem well, and if they know what they're going to be going to the dash to find, let them skip that step. In case something new happens, they can set up a new alert/workflow/whatever for the next time it happens.
This is part of the expectation of business analysts, that they will surface and communicate data not immediately being asked for by executives or other stakeholders. This usually means running a meeting and presenting findings, and then pairing with another stakeholder to flesh out further opportunity.
If there's no one to pick up the baton, there's very little that can be done about that.
* Provide a clean minimal view of the data that other teams can pivot table on top of or download in the BI tool. We do onboarding, support and all that for this data.
* Provide dashboards and reports for more complicated ongoing questions that require data that's not cleaned yet.
We don't use Looker but based on the sales demo it seems optimized for this type of workflow. The core analytics team maintains the data and complex workflows while other teams have business analysts for day to day questions.
With big data warehouse, more often than not, the fancy dashboards would not run fast enough without a dedicate pipeline that materialize data before hand. Too many useless dashboards and your data team would not have enough manpower to do anything else other than operating the existing pipelines. As a result, new dashboards become slower than old dashboards due to lack of maintenance/optimization, and everyone complain about the poor performance.
A lot of what users need can be done this way in my experience. Also giving their people access to customise reports/exports can go even further. I'd advise against direct access to “real” production though: have the dashboard and custom reports run against a (probably readonly) replica of some sort¹ so runaway reports don't impact live app performance.
[1] via replication, log shipping, etc... — all major databases should support at least one method that can achieve this with minimal latency in updating the replica under normal conditions
In my experience, no, as the company gets to any complexity and size. The production DB is often optimized for use by the ORM/language of choice or there's twenty of them for microservices. This makes it really awkward to do some things in SQL. So now your engineering team is getting requests to keep changing the prod DB schema so analytics can be done more easily. More over the prod DB is likely 10% of the data your business needs. There's a reason companies like Fivetran exist to slurp fifty data sources into a warehouse. Marketing wants google and facebook and mailchimp data. Operations wants saleforce data. Finance wants stripe data. And they want all of this tied to the inhouse data.
Sure you can make reports in each of those services but then you need to tie them together and then people complain of differences in metrics and then you have a bad time.
edit: Not to mention analytics on what users do. Always fun when there's a billion row table in your production db for that.
I was generally able to create better results via power bi, ssrs, or shiny
But most people seem to log in once or twice and then never again. And if you ask why it’s often because they couldn’t find the answer to their incredibly specific one-time only question, so they decided it’s all useless and they continue to work on their own personal homegrown BI-suite in excel.
Also, we use Qlik and I absolutely hate it. Luckily I don’t have to work with it too often but whenever I do I feel like I’m always fighting it. Using Power BI and Tableau felt much more like working together to solve issues.
Hah this is probably me. The university I work for uses a BI tool (MicroStrategy) to track students/majors/etc. But I usually find it easier to use the BI tool primarily as a way to get a CSV or XLS export of enough source data to answer my question, and then do the actual analysis in Python+pandas. I can probably formulate my query in their weird proprietary query builder, but it seems unnecessarily complicated and I already know how to do data analysis in Python.
If you're talking about Power Query, the best way to use it is to write code directly once you observe what the interface generates. It's a rather nice kind-of-functional language that really comes into its own when you start using/creating higher order functions.
Power Query gave me an intense craving for a form of SQL with a similar facility for functions.
For example, I have a lot of experience writing Oracle SQL, and Power Query didn't offer "grouping sets". But I realized it could be implemented using functions. It would be so great if SQL supported that.
Why aren't you providing your dashboards via Excel, then? What exactly does the fancy BI software do that Excel won't? I have used Qlikview but not Tableau or Informatica, and I'm a bit fuzzy on what people mean when they say "Power BI". Is using Power Query what you would call Power BI? Or does that mean using the data model stuff? Power Query is the part of Excel that I particularly like because data kind of stays where you put it.
I've used both DAX and M (Power Query) via recent versions of Excel. Obviously there are plenty of ways to interface with data.
https://en.wikipedia.org/wiki/Power_Pivot says both Power Query and Power Pivot are part of Excel.
What I'm seeing there looks as though it is targeted at people who make the decision to buy products like this, but not to use them. I'm not seeing what couldn't be done with Excel 365 alone. There are screenshots with verbiage about how well it works with Excel. And they look like...Excel.
I guess it supports R, and maybe makes it easier to seamlessly refresh a report. That's aggravating though, that it makes me suspect Excel might be purposely rough around the edges in order to not displace Power BI.
As I've said a few times, I like M and wish they'd commit to it.
I honestly suspect Microsoft is repackaging functionality that already existed for many years in their "ball of mud" software, with a shiny new wrapper and better defaults, but practically unusable, because people who license software are easily manipulated.
Power Automate is the most awful example of this I've used so far.
You say it's a competitor to Tableau, well, Tableau is a competitor to Qlikview, and I've used Qlikview and Excel enough to consider Excel better in every way that I care about. (My organization uses SharePoint, so that's how reports are usually distributed)
Power Query is the only part of this that has seemed useful to me.
The Data Model and DAX in particular seemed to me like distractions that I wasted time on before discovering the M language.
Saying Power Query is mostly unrelated demonstrates that the confusion is not all mine.
So then we have to get our db people in to rewrite all of their reports so they don't do this which eliminated the benefit of having the customers do it.
In the end there is just no avoiding having experts build the dashboards.
Interesting as we're looking at going long with PowerBI embedded for a client facing app and ditching our home grown effort which is always playing catch-up and costs a fortune.
When it comes to business stakeholders, the biggest obstacle is trust. If you're not a data person, or a developer, decision logic being gated behind code is a scary black box. "If I can't control it, I can't trust it". That feeling gets even worse when we build systems that are controlled by AI or ML.
I think we need more solutions for data teams that allow business stakeholders to take part in the automated workflow deployment. Let them see how things are connected together. Let them verify that the decisions are being made correctly every day. Let them tweak levers so they have a say in how things are working. That's the only way to move beyond the current environment where everyone wants a dashboard, but no one looks at it.
That's a big problem I'm aiming to solve right now.
The technology is getting good enough with cloud warehouses that most of the logic can be defined in SQL, with a script only needing to map the results back to an API.
Current domain is SaaS. I've found many of the same types of opportunities on the sales/marketing side of things, but not at quite the same scale.
Agreed that automation it will always take more time than a visualization. If we can shift the conversation towards the expected action that will be taken off the dashboard, we can hopefully save thousands of wasted hours on visualizations that don't get used, redirecting those hours towards more impactful efforts.
I would love to expand my mindset and hear of domains where the complexity is far higher. Feel free to shoot me an email.
I think the visualisation matters though. It's much easier to convince people that the automated actions are worth doing once the related data is visualised. And at that point, whatever provides you both may be the best solution.
Is the platform still easily used by business users? Or does it primarily become a gated way for Data Engineers to build/action on data?
A lot of decisions can't be easily visualized in your standard dashboard. You can forecast the potential of those automated decisions, but that's still black box logic. You really need more of a middle ground that lets you preview "given these inputs, show me what the output would be".
I don't know and honestly I don't believe there's a system that actually achieves it. But this it does go beyond BI in actually being aimed towards running code based on results. Which makes it closer to Labview-with-graphs than to BI.
Essentially, I first implement the engine with the output being an automated list of 1-click actions (typically a set of links on an email, but could easily be a dashboard button or anything else), so the system default is "detect, but don't do". After an evaluation period of the system actions, we move onto the second part, in which the system performs the actions and produces a log of activity, plus a list of 1-click UNDO actions to the user.
The idea is that a) the system earns trust from the stakeholders due to their direct involvement, b) the system can enjoy some supervised learning from someone other than the dev team and c) worst case, if it never earns enough trust, I've still saved a stakeholder dozens of hours of work sifting through dashboards as opposed to taking a look at a pre-filtered list. A few systems have turned out well enough that they never move from Phase1 while still being considered a massive win.
Is this all internally built?
Yes. A lot of darts are being thrown as to where XR fits on which Hypecycles, but one of the most promising uses to me is as a new enduser engagement layer for BI(g) data. The new UI/UX possibilities of immersive 3D have a low level effect on what/how information can be rendered in meaningful ways. It's still early days in this space, but here is a quick list of some work in progress for anyone interested:
[1] https://flowimmersive.com/ Focused on Data driven story-telling through an inhouse 3D visualization tool. Strong on AR collaboration and responsible for "The Data Guy" popular on TikTok.
[2] https://www.badvr.com/ Industry focused on using XR for specific applications like 5G Radio Coverage heatmap modeling and Smart City interfacing tools.
[3] https://3data.io/ Use of XR for IT Operations Center applications (NOC/SOC)
[4] https://github.com/ACEMS/r2vr This project lets you output basic WebXR 3D visualizations using R
Allowing stakeholders some leeway to conduct their own independent analysis (after a short training session) has allowed our strained data team to hand over simpler analysis to focus on the harder problems.
Less is more.
The goal of properly exploiting BI comes with many prerequisites which sound superficially reasonable but turn out to be decade-long side-quests. Things like having all your data in one data model. Things like understanding where your data comes from, and exactly what it means.
These prerequisites are easy for small orgs, but small orgs benefit least from BI and typically get better bang-for-buck from Excel.
Large orgs find themselves mired in the political meta-problems of meeting those prerequisites, like joining up the fiefdoms that own data sources with the cabals that run data governance and the accountants who want a return on the investment of simplifying a sprawling legacy estate.
BI tools are generally incredibly poor at dealing with the bag-of-spanners heterogeneous data landscapes that exist in these organizations, and their analytics nirvana remains largely unattainable.
The trend that the article describes - towards simpler, composable BI components, each with more modest goals - is progress. It helps move focus away from the relatively-easy problem of visualization, to the rest of the data stack.
It's not that they are poor, it's that they eventually require some moderately-challenging coding - because there is only so much complexity you can hide behind point&click before things become untractable. Programmers get bored doing BI plumbing, and those commercial ecosystems are quite secretive and often expensive to train in; so the talent pool is small and costs can skyrocket pretty quickly, making the whole proposition unappealing.
Maybe if we accepted that "BI Programmer" is a respectable career and something that is fundamentally unavoidable, providing clear career paths inside companies for it, instead of trying so hard (and failing) to get rid of such figures as soon as they progress a bit, there would be more predictability and less angst in the space.
Or did you mean, "I have no idea what business needs all this data for", a much broader question?
Today you also have Pulsar, Redpanda, Kinesis, PubSub, and NATS off the top of my head. You can do the same streaming patterns with all of them, with the exception of log compaction probably. It's an architecture so amazing that the world needs 6 redundant implementations. I think my OP implied something regarding that.
https://github.com/vectorizedio/redpanda/tree/dev/src/v/storage - see all 'compaction_*' for deetsLinkedin saying Kafka is good for them is not saying much.
Must be solving some problem for them, eh?
I share the question about a solution searching for a problem. I will say that I continue to see efforts being made that seem focused on adding a bullet point to some executive leader's yearly accomplishments rather than specifically providing value back to an org . . .
Like why have databricks, azure synapse, AND snowflake on the same picture. I’m certain I could do everything they say they’re going to do with this shit with an integration tool, snowflake, a machine to run python on, and our BI tool in the same cloud.
This hits painfully close to home. Our org just "finished" such a build, mainly done by a large contractor, and the result is a messy nightmare.
> integration tool, snowflake, a machine to run python on, and our BI tool
Yep, my next campaign will be to gradually pare it down to this. Snowflake is wonderfully capable.
I've been working in the BI industry for almost 20 years and I don't know why the author believes that the "Original BI" includes ETL/storage. He mentions BusinessObjects and Microstrategy as an example of traditional BI tools popular 20 years ago, but neither of them had ETL/storage capabilities. Both were basically visual SQL generators. (Although, BusinessObjects later acquired an ETL company). Qlik has a bit of everything - ETL, columnar storage, and dataviz. But that's more of an exception than a rule.
Nevertheless, the author has a point - BI seems to not have a clear path forward. There are experiments with natural language queries (won't work), automatic text summary generation (useless), AI-assisted automatic insights (might be a nice feature, but barely a "product"), and a few more. The self-service story of BI never really took off (I believe self-service ETL is a more interesting story [1]). Neither did story-telling. Dashboards are oversold and are mainly used to impress CxOs to close big deals. Geospatial dataviz is useful, but has limited application.
I would envision two possible directions to advance for the BI industry:
1) Analytical notebooks, something like Jupiter Notebooks but no-code and for BI might be a good idea to explore, but I haven't seen anything like that.
2) No-code/lo-code analytical app building. Reporting and dashboarding tools are just highly specialized app builders. Why not make a step further and generalize it a bit?
[1] https://easymorph.com (Disclaimer: it's my company)
I've done some BI, but mostly DW/ETL, for over twenty years... And I can't remember ever seeing ETL viewed as a sub-category of BI in any project I've been part of. It hasn't even been mentioned as such -- except in weird articles like this -- this side of the turn of the century. Before that, yes, the categories were muddled. But not since then.
> 2) No-code/lo-code analytical app building. Reporting and dashboarding tools are just highly specialized app builders. Why not make a step further and generalize it a bit?
Sure. I've long thought the best tool, one that could do both ETL and BI -- is an IDE with a full-blown programming language and a good library: Something like D̵e̵l̵p̵h̵i̵ Free Pascal / Lazarus.
Yes exactly! Dashboarding tools only get you halfway - there's no easy way to take action from within the dashboard itself. I see analytic app building as a cross between Tableau and Airplane/Retool
IMX, the core problem is that all BI software is built around the idea that everything is a financial value that can always be arbitrarily aggregated meaningfully, and that looking pretty is more important than presenting the data in a way that your organization actually finds logical.
It's a way to sell very expensive software to decision-makers, while also producing software complex enough that it can't be configured well enough to properly evaluate until well after you've already bought it. Only then do you find that the feature set is very broad, but very shallow. It's only 18 months later that you discover how limited the report writer is, or that you have to do it this one way for everything even if you really could use it formatted slightly differently.
It's like using a pivot table in Excel and trying to control the order of the columns, or to make one table aggregate two values distinctly, etc. You end up with 10 seconds to pull your data and create the table, and 4 hours trying to get it to display in the manner you want before giving up.
BI, like every other tool doesn't exist in a vacuum. It has to be present in an environment that has peopeople that know how to help themselves use the self-service attributes (i.e not Boomers), it also has to occur in a culture where KPIs and metrics are understood, and used in day-to-day discussions of the business, there should also be a training component to deploying your BI solution too.
Without these aspects being present, your BI solution will devolve into either a sales prop with pretty colors, or it will become a web version of Excel.
I used to work directly modeling data for various sized companies on behalf of two different BI sellers.
As I see it there are perpetually two problems:
1. BI is marketed at and sold to anyone needing any kind of data visualization capabilities, and most packages I've worked with can do this but that's not really where they shine.
2. They really shine when you have huge amounts of data stored in different systems and you are trying to build an environment up where you can coalesce that data into one place and then visualize it, and you need to routinely report on these sorts of things.
I think most BI solutions fall into the same space as JIRA does -- highly-customizable solutions that are sold as "Turn-key" without any warning ahead of time that it'll require someone (or several people) on your team become ridiculous levels of expert in areas that don't necessarily help the business.
It's a specialized skill most places don't really need when what they're looking for is a simple reporting package that can connect to just any database.
Similarly, I saw a lot of people with BI insisting they had "Big Data" and thus "needed" things like Hadoop in order to process their stuff, when in reality the MySQL, SQL Server, or Oracle DB they were already invested in would do the same job faster 99.9999% of the time.
I worked with consumer apps and it was used for making business decisions related to return of investment and marketing.
But it took some time to get there and before that I think most decisions were based on questionable assumptions from half-baked results.
What was the final building bricks were to be able to calculate the lifetime value of recently acquired users, but also to collect all the types of revenue and cost and through various tricks (based on user base numbers mostly) break it down into countries and marketing channels.
(Though even in the target audience, I'm not taking much away from this. Tools should be better. Yep.)
Sure, each of these tools are better suited to accomplishing a very specific task with your data. But the problem is that now these tools aren't talking to each other. No one has a single place to view their data end-to-end. No one can show me every automated process in their organization that relies on a specific table. Everything is tied together with trust. Trust that the Data Engineering team won't deploy bad code. Trust that the API won't change how it returns data. Trust that the data actually loaded completely.
This trust-based system means that one problem during ingestion/cleaning ends up spelling disaster for all of the downstream ML models, reports, data extracts, etc. and causes a day worth of headaches to resolve.
If BI used to hold everything together under one roof, we need something to make these tools talk to each other. The only solution, as I see it, is to build out better data orchestration that effectively glues these tools together and lets you see the "big picture" with your data.
I don't think BI analysts have to get worried quite yet.
But, come to look at it in hindsight, my comment may have been misplaced -- maybe it should have gone as a reply to the same comment you replied to, with its "It really is amazing how far you can get..." At least if they meant "amazing" in the sense of "surprising". To me, what's amazing and surprising is only how many people don't already know this.
The main goals are:
- Easy accessible for novice users. We want to make the Google for BI to help empower non-technical people (no more salespeople asking for reports from devs)
- More advanced editors for the more technical people
- Advanced alerts + integrations to 3rd parties
- Later on proactive reports
Hit me up on @philipanderse if you want to test it out.
Maybe I am a naive fool but IMHO the benifits of BI are reconciliation of the business model:
- reconciling processes - reconciling data - reconciling governance
Its a painful process to impliment and you will hit all of your compliance, governance and change management issues head on.
IMHO if you accept that this is what you are doing then its potentially very valuable for your business, but it must be done humbly with humilty as every single measure and many dimensions are each a project on their own.
The pretty graphs are largely irrelevant apart from making employees feel valued and to re-enforce abstractions via visual metaphors.
I get pissed off when end users are not trained in consuming the database warehouse through tools like excel or access. This project should include general IT education programme for most companies.
Also.. if the end users cannot use a pivot table then you need to teach them that and prototype a bunch of stuff in excel before you go anywhere near BI of any flavour.
All the while things like Google Docs and Slack made collaboration around documents and ideas much easier with @mentions, threads, etc.
So BI can be a lot more useful if it is 1\ accessible anywhere (not just desktop or a crappy mobile app) 2\ collaborative -- bringing the mentions, shared questions, and ability to make decisions together 3\ not crazy expensive (looking at your tableau + looker)...because at these high price points the tool ends up not getting offered to everybody, it ends up being more limited in who gets access and hence less useful across a company
Full disclosure, I built Zing Data which is an app for mobile first business intelligence and is free for small teams. Works with PostgreSQL, MySQL, and Google BigQuery. https://www.getzingdata.com/
Would love any feedback folks have -- we're actively improving it and I'm sure this community has a lot of great ideas!
Just like bricklayers haven't even remotely been replaced with robots, I doubt the data-mungers will be, either. We are probably on the cusp of a whole new generation of "data people" who will have careers that span a generation, and do nothing but sift through data.
Because I believe that BI should be done by people who have a mixed set of skills between stats and programming.
A programmer alone can’t make the stats, and a mathematician can’t build a reporting tool.
Jig is a shortcut to build reports quickly, but I am somehow convinced by javascript is just not good enough to do that job.
I think I might have to try and design a language that has indexed containers as first class objects…
Even as a software engineer doing basic analytics charts for the service I own I don't want to be writing JavaScript. Ideally I wouldn't be writing SQL either.
And companies are trying. Ex:https://www.bobsguide.com/articles/barclays-gordon-risk-mana...
But I also see spending more on these tools, in the name of innovation, because some bigshot likes tool X and that's what he wants to use. And guess what, now you also need it to be made available in the cloud.
Tableau is a nightmare (no Linux support for a start, nevermind the hassle of editing each dashboard individually).
I wonder how long this took to write.
we use Microsoft PowerBI. it work great within the Microsoft ecosystem. if you try to pull data from NetSuite with PowerBI. it went to crap.
If you and I sit in a room and race to analyze a few CSVs of data, I bet I can find compelling trends faster using a BI tool.
TL;DR: I believe the trend will be to have an ecosystem-like approach to BI with tools complementing each other and communicating with each other through open standards.