Data engineering learning path with recommended resources
awesomedataengineering.com
awesomedataengineering.com
As someone who has worked for the past several years in this space, I'd say the biggest problems in data engineering are wholistic in nature. Sure, you need to know Python, SQL, Data Warehouses, Data Modeling, etc., but to me by far the biggest problems have to do with the entire architecture i.e. How do you extract data from potentially unreliable data sources, pull that data into some staging area, build further workflows that base off of this raw data to reliably update or create data warehouses/marts or deploy ml models. How do you allow everyone in your company to access and work with the data in a compliant and secure way? How do you test any of this? How can distributed teams, sometimes technical, sometimes more business oriented interact with the architecture and add/control data and release it into the overall company data stream? Has anyone found a reliable and maintainable way to setup CI/CD for company data architecture/pipelines/projects?
To me these are the big problems. And if anyone has any resources for any of these topics I would be super interested, since I deal with these problems daily :)
The orchestration part handles the workflows that comprise your ingestion and ETL processes. These are like managed cron jobs specific to data engineering lifecycles. The bespoke part of the architecture is what you'd compose together to handle all of the other requirements; for example, what applications do you build, and how do you design your data warehouse, such that the architecture can be used by both data science and marketing teams?
I absolutely agree. The hard problems are either organisational (how to communicate to analyst and have agreed work method with business?) as well as dealing with third party unreliable resources.
I feel like these things you can learn only through experience. No written resource can reliably transfer this knowledge.
I think data engineering especially is something that requires at least apprenticeship to get into. Both for juniors and for senior developers transferring to data engineering position.
data engineers are dependent on software engineers. and the software engineering is the more difficult part IME.
- How do you make sure the users are not writing badly optimized tables/pipelines that end up consuming too many resources.
- How do you facilitate data discoverability so they don't end up creating a new table where 90%+ of the data is already present in another already existing one?
- How do you make sure they are mindful with the way the model the tables such there are not too many files due to bad partition/bucketing, compression is leveraged and good datatypes are picked?
The worst part is, these things tend to grow organically where the original ancestor of everything is engineer #3 of the 5-person startup who decided to write a cronjob to dump the prod db every night to feather files on a samba share. Then one cron job becomes three. Then ten. Then you're using S3. Then there's dependencies. Then you're using Luigi/Airflow. Then you're using Spark. Then you're constantly messing with partitioning and performance. Then you're using Hadoop and YARN and configuring queues and capacity scaling. At this point the "infra" team is 10 people and there's 20 data scientists. Then you triple in size a couple dozen times and now it's time to figure out how to retroactively slap security controls on top of everything.
That turned into a bit of a rant. Honestly I love this space but I agree, as hard as the software bits are, the hard part is the entire picture as a whole.
an experienced databricks/snowflake architect (or a couple of them) can easily set up and maintain data lake that supports most of what was mentioned.
Overall I absolutely agree, that rather than learning python or SQL, it is much better use of one's time to learn/get certified as Data Lake Architect and be able to create a large data lake from scratch and set up pipelines and maintain them.
* https://greenteapress.com/wp/think-python-2e/
* https://automatetheboringstuff.com/2e/
* https://dabeaz-course.github.io/practical-python/Notes/Conte...
Also, I'd highly discourage tutorialspoint as a resource. Here's an example of them rewording another tutorial as their own: https://twitter.com/nixcraft/status/998248317661335552
It’s my favorite role I’ve had to date and I’m really happy in it.
As for the SW at the big firms I do not know. I was hoping to end up at an Apple or Amazon or Google so the culture or stress being not so great is disconcerting.
I really appreciate the effort but as an anxious person I always paralyzed or disheartened by the road ahead.
For instance, one of the recommendations is Learning Python, 5th Edition - Mark Lutz. This is book alone is a tome.
But anyways, it looks very well presented. Much better than plain bullet points. Well done!
Interesting rendering as far as webpage is concerned, what framework did you use? Maybe a tutorial on rendering Json data to a webpage like this is really helpful.
Good resource.
I’d encourage folks to think about their data retention policies early. Build your data architecture with privacy in mind. Regulations like GDPR can require you serve customers a copy or their data or delete their data from your systems.
Don’t store things you don’t need. Keep retention policies. Please protect customer data.