The Next Big Challenge for Data Is Organizational
locallyoptimistic.com
locallyoptimistic.com
I've been trying to populate the catalog with a bit more explanation of what things are for + data dictionaries, but it feels like boiling the ocean.
This means someone may need access to raw data, data after it's been cleaned and put in the data lake, data in an intermediate state at some point in the transformation chain, pristine fully joined master records for some data in the warehouse. Some teams shouldn't have access to certain types of data, or data earlier and later in the stack.
There's not a central way to do access control at a fine grained (row in some cases, files in others etc) level across layers and multiple services/dbs in the data stack. At least no plug and play way from what I can see. I'll be in the midst of building that layer for an entire enterprise organization soon and I expect it to be enlightening.
I'm at the same time putting in some leg work to explore how common this problem is across other companies right now for obvious reasons.
There are lots of good tools for data profiling, but a good data catalog is something that is structured, easy to use and actually separate to the data. By keeping your catalog separate to the data you get more people involved which means more people can find and know who to talk to, to take advantage of data.
That’s why we’ve built our startup - aristotlemetadata.com - to give organisations a tool for users to collaboratively document why their data is important and what it means conceptually, along with raw column information (which is an easier, solved problem).
Building data culture is a bit more challenging, but again tools play a role, as if they can see others getting involved in data governance, can see examples and see others in their organisation getting more value from data more efficiently they are more likely to want to be involved too.
notepad.exe is enought. The main "unsolved" problem is keeping track of the documentation, which needs secretaries on non-management level which for some reason is a no-no.
There seems to be a lot of hype and hope to organize domains of data much like you would products of a business. I just don't see it happening today, largely because people in engineering/product, upsteam of DE teams, are unwilling to put in the effort.
That said, check out https://datahubproject.io/ . It's the closest OSS I've found that is following the spirit of Data Mesh.
That's an older challenge.