We need new data books, so we started one: Cloud Data Management
chartio.com
chartio.com
I could just as easily ask the same about databases, or dozens of other tools in our toolchain. Makes me wonder if there's space for something like SQLite and testable.
Shoot I must be getting old, rambling on a tangent. Get off my lawn! ;)
Because Tasks are just functions you could test your tasks with regular unit testing and mocking methods.
But I agree a lot of data people seem to think interactive is good enough.
Heres a rewrite of the airflow tutorial in prefect: https://docs.prefect.io/core/examples/airflow_tutorial_dag.h...
My advice is try it out and see if you like the api. I think that the airflow ui is still a killer feature in 2019 though.
Theres also this post from the creator of prefect on why not airflow: https://medium.com/the-prefect-blog/why-not-airflow-4cfa4232...
I haven’t tried dagster yet, how does it compare to airflow?
Possibly a side effect of nosql movement/proliferation that these basic things became esoteric. Like a vet care for dinosaurs :) The tooling around data pipelines, like debugging, profiling, lineage, etc was too a typical stuff of the ETL tools of yesteryear like Informatica for example.
When I write pipelines I write unit tests for the business logic and integration tests for the overall process.
Whether you use SSIS, Alteryx, Spark whatever you can write tests and you can debug using known inputs, rather than running an entire pipeline to see that is works!
These are my typical thoughts on ETL testing:
https://the.agilesql.club/2019/07/how-do-we-test-etl-pipelin...
https://the.agilesql.club/2019/08/how-do-we-prove-our-etl-pr...
https://the.agilesql.club/2019/09/how-do-test-the-upstream-d...
https://the.agilesql.club/2019/10/how-to-test-etl-processes-...
ed
Great work by the guys from chartio!
Besides the need for a new data book, we realized that it needed to be of a different format, as the space is moving so fast the expertise is very distributed.
Tech/web companies deal with massive amounts of unstructured/semi-structured data ingested at a fast clip, so the architectural thinking here works.
However I would argue that many traditional enterprises whose major sources of data are primarily highly structured (SQL databases), a lake is actually not needed.
A young data engineer working in a traditional enterprise, enamored with the idea of data lakes say, might try to ETL SQL databases into an object store (add RBAC etc.) only to rebuild it back out into a data warehouse. This will almost always turn out to be the wrong approach.
The simpler and more manageable approach is actually to federate existing databases, add cataloging etc. and not even use a data lake at all.
Which is why I mentioned federation. Most data sources in an enterprise are SQL native or accessible and it does not make sense to dump a SQL database into S3 just to be able to combine it with other forms of data.
Federation means you can operate across multiple databases (eg do Cross database joins etc.)
I tend to joke that I'm a digital janitor and now I'm a big data janitor. Well, things are going well enough, but cleanup on a live system with users and zillions of reports is incredibly difficult. Part of the issue is isolation and size: too many have too much access to too much data. It's overwhelming and it also leads to a lot of high cost because queries aren't understood and they're run against massive datasets.
I've been educating myself as well as I can between fixing errors and this book is just the thing I need to calm my nerves. It totally makes sense how the stages are set and even just this blog post overview has given me ideas of how to carve some things up.
Kudos, cheers, and all that. Happy Halloween!
So, perfect timing! Really looking forward to checking this out.
In the book we use much of the old terms and recommendations. Most of the high level organizing is still totally right - but a lot of the optimizing and work done for performance and cost reasons is very different now.
For example ELT makes now much more sense than ETL for the reasons Kostas wrote about here: https://dataschool.com/data-governance/etl-vs-elt/
And many things previously done for cost and performance reasons are just not relevant anymore thanks to the big innovations in C-Store warehouses.
The difference now is, you have Hadoop and cloud providers that will take credit cards and give you as much space as you can pay for. The concept is not new, it was just cost was a factor back then because capacity was fixed and memory was expensive.
the only thing that has changed is the commoditization of hardware has allowed for different behaviors that would have been cost prohibitive.
It allows you to not do E & T & L all together. It's really nice (less complex, easier to implement, less costly, and more flexible) to have that T part pulled out and done after.
The T being after the L means you can do that stage more simply in just SQL (with views or materialized views - possibly with the help of DBT), as opposed to some vendor interface, or python/R/etc script.
It also means that rebuilding your warehouse is much less significant of an ordeal. When the structure of source data changes or if you want to make some migration of the schemas of the warehouse you don't need to also re-run your ETL jobs and start over from scratch.
C-store doesn't solve for denormalization or provide any advantages when you're copying data all over the place to avoid joins.
And though ELT is a very common standard now - I don't know of a single book that recommends it or explains why that change has happened. Just one of the reasons for writing a new data book.
2. C-Store does largely solve pervious performance and cost issues with denormalization. We also write a bit in this book (more to come) on how to avoid doing all that copying.
There are a number of people starting to talk about Star Schemas having little gain on modern stacks. The performance and costs gains are automatically done now with C-Store warehouses. Fivetran has a great post on this https://fivetran.com/blog/obt-star-schema
in the world of fixed assets (on-premise data warehouses) this was a concern, I'm curious to see how this plays out in these Cloud Providers because to me it seems like they will be more than happy to rent you as much space as you want and everyone is happy until the bill comes due.
Doing materialized, flat tables everywhere is great for reporting performance but the tables will not be updated as quickly, there will be redundancy in storage and it will be difficult to sync time dependent dimensions.
We try to explain some of that here - https://dataschool.com/data-modeling-101/row-vs-column-orien...
And Fivetran did a great benchmarking of it here - https://fivetran.com/blog/obt-star-schema
The architecture of the C-Store warehouses often removes the benefits of materialized views. This is why for a very long time Redshift didn't even support them - they insisted they weren't needed as they didn't improve performance over regular redshift significantly.
would have loved to have this book when we were building this at my previous co.
very approachable in how it explains each segment on the whole and zooms in on them individually.
would have saved our team a lot of time as we came to very similar conclusions but over the course of a few months
As an example, doing ETL with drag n drop tools, and not in code, is a dying skill in the industry.
[edit] Basically I'm wary of companies commenting on data standards, where those same companies also sell a product in the data industry. You're probably going to find more honesty on guidelines from open source contributors like LinkedIn, Netflix, Sitchfix, etc.