HNHacker News
TopNewBestAskShowJobs

verbbis

1 karma · joined April 16, 2021

submissionscomments
verbbis··on What the Heck is a Data Mesh?
"Trino pushes the queries down to both systems and processes the rest in flight."

This is how many data virtualization techniques work as well (e.g. PolyBase from Microsoft) - predicate pushdown is not exactly new. That is why I mentioned both approaches. However, according to my experience, the degree to which this helps is highly workload-dependent. Do you happen to have some reference which would help me understand how e.g. cross-domain joins between large datasets can be optimized with this approach?

"Are you saying teams aren't allowed to store immutable data in PostgreSQL?"

Of course they can. But unless you assume everyone (incl. the numerous COTS and legacy applications in any "normal" enterprise) is practicing event sourcing and/or their internal data models are somehow inherently understandable and usable by downstream consumers, a pipeline of some sort is required. That was my point. And if so, what's the difference for the team to target a dedicated DB within a cloud DW instance that speaks the PostgreSQL dialect - or close vs. a separate PostgreSQL instance?

Furthermore, if you think standardization of tools is not possible within an enterprise and everyone just does their own thing anyway - and mind you, despite suggesting Trino as one such tool yourself, I have low hopes of getting the data mesh standards for governance ending up being adopted either.

verbbis··on What the Heck is a Data Mesh?
I have to disagree with almost all of the points you raise.

The degree to which you wish to (re)model data is a design decision - also within a DW. This is what I mean by divorcing the tech from the practice. Methods to store and manipulate semi-structured data exist within cloud data warehouses and there is nothing inherent in the technology preventing from extending said support. Also, it seems to me, that even in advanced analytics incl. ML one eventually works with the data in a tabular form anyway.

When it comes to performance, I was referring primarily to e.g. cross-domain queries. That continues to be challenging in data virtualization / federated query engines.

You (and data mesh) focus on the domains - and nothing prevents carving out ample portions of storage and compute from a cloud DW to them and doing the exact things you propose. Except that "federated computational governance" and enforcing standards is now heaps easier since everyone is relying on the same substrate.

Data mesh targets analytical workloads - and surely you do not suggest e.g. hooking Trino up directly to operational, OLTP databases? One has to, as least how things stand today, copy the data somewhere anyway for not the least historical analysis, as well as often transform it to be understandable to a downstream (data product) consumer. So you will have "pipelines" no matter what - even in a data mesh.

Not only are the domain teams free to build the aforementioned pipelines and model data in any way they see fit - within a cloud-native DW, but also the DataOps tooling available in this area is already relatively mature.

The only part I can somewhat relate to is about "forcing" to rely on a specific piece of technology. But I think that is just something one needs to accept in a corporate setting and a balancing act. And I'm talking about the majority of the regular companies out there, not FAANG. Also, there other arguments to be made - such as avoiding vendor lock-in or if the capabilities of the DW simply does not cater to the specific problem your domain has. But these are not the arguments you made.

verbbis··on What the Heck is a Data Mesh?
Few data mesh proponents ”force” a particular storage medium - and the concept is largely agnostic regarding to this. But lots of early implementations in the wild have decided to standardize on it - either on some cloud object storage or, indeed, a cloud DW.

One cannot argue how much it simplifies things in terms of manageability, access, cataloguing, performance… in an already complex architecture. Especially since no reference implementations exist.

I understand that if your persistence layer is heterogeneous from the get go, layering on top of it might be a solution. But it is also an additional layer that needs to be managed.

Conversely, in your opinion, what would be the shortcomings on centralizing on a modern, cloud-native data warehouse (tech, not the practice)? I see this being articulated less often.