Data Engineering Cookbook
github.com
github.com
Computer science is: 1) Daunting to beginners 2) Changing constantly
It is absolutely a field where people suffer from "I don't know what I don't know". 100-ft overviews of best practices and new buzz-words are refreshing, in my opinion.
The criticisms I have of this document are all preference. I do not want to watch the podcasts or youtube videos, especially if I just want a high level overview of something. A ctrl-f of 'andreaskayy' returns 12 results. This guy's self promotions is everywhere, which is fine for some people, but it makes me think I'm getting 1 man's opinion on everything, and not an unbiased explanation of different technologies/methods.
I think fundamentals of CS are crystal clear and are not changing constantly at all. There are ideas, techniques and principles that were set decades ago and are still fundamental. Things like "program is data", certain set of abstract data structures (array, list, map etc), certain algorithms (BFS, DFS, topsort, sort, tree-search etc), certain mathematical methods (proof by induction, least squares, SVD etc), basic software engineering principles (abstraction barrier, abstracting something out, data types, parallelization, runtime polymorphism etc), and system concepts (OS, syscalls, cache, filesystem, database, branch prediction etc) are always there and are very much sufficient for most people to read new ideas on CS. If you have an undergrad degree in CS and studied most of these things and did well, you can pull up any blog post/paper and with sufficient amount of effort, you'll be able to learn new things.
CS suffers from a lot of hype. We have a lot novel ideas and they are actually very useful to solve larger problems. This gives the illusion that CS is constantly moving and changing. Whereas, the battle-tested ideas are there and the newer ones are extension of these ideas. My opinion is once these new ideas can be simply explained using CS-primitives I listed above so that someone with BS in CS can understand, they'll be actually widely useful topics. Until then, they're concepts experts playing around with and seeing if they're effective at solving some problems.
I agree it is. But just I took a quick look at the book, while I understand the book is not done yet, looking at the table of contents, I doubt this would help. Like, do you have to cover kerberos, IP Subnetting, OSI IP model, Agile, git, REST API, docker, and many more in a single book about "data engineering"? If anything, it would confuse beginners even more. Its like the author tried to cram as many buzzwords as possible into a single book.
However, I don't think the purpose of this book is to cover these subjects in their entirety. Most of the time, with books like these, I would just like to see definitions, use cases, why it's used, maybe even what it replaced in the space it exists in. Materials like this aren't meant to make you an expert, they are just meant to show you what's out there, why it's out there, and (maybe) what was out there before. They give you context with which to google things.
Either way, everyone expects something different from educational materials. I personally would not use this book, but I get the idea he was going for.
1. Given the title, I was expecting a more traditional Cookbook style book. I.e., "How do I..." followed by one or more recipes to answer the question.
2. The blueprint on page 39 is a good start: it includes the 4 main processes for data engineering. However, Display is only one use case. Training models, for example, is another. It also ignores real-time vs. batch processing. This comes up later in the book, but could be diagrammed more clearly. There are a lot of recipes for the overall architecture, and for each subprocess.
Data Warehouse vs Data Lake chapter is a single podcast link. The Hadoop chapter is 5 pages, mostly used by large diagrams and the docker chapter is less than 4 pages with half of the sections empty except a heading. The REST API chapter is less than 2 pages with a blank section headed OAuth Security.
Data Visualization is entirely blank. The database chapter is mostly empty except for text about HDFS, and just links on MongoDB, ElasticSearch and InfluxDB. Apache Kafka gets its own mostly blank chapter.
Most of the beginning chapters seem unrelated to data engineering. 3) Learn to Code. 4) Getting Started with Git. 5) Agile Development. 6) Learn how a computer works (section 1 is subtitled "CPU,RAM,GPU,HDD" but the chapter is empty). 7) Computer Networking - Data Transmission.
How do you think a book gets written? Obviously you don't think that someone sits down, puts finger to keyboard, and then a book bursts into fruition. This is a work-in-progress kindly made freely available. Is it really fair to criticize the author for not having finished it yet?
So I need to have written a book to be able to download a PDF and see 85/100 pages are blank? I work as a data engineer and can tell you 50% of these chapter topics are not directly related to data engineering.
There are no chapters in this book even close to 10% finished. If you want a book recommendation I'm seconding the suggestion in this thread of Designing Data-Intensive Applications. I have a copy 3 feet from me at the moment.
> This is a work-in-progress kindly made freely available. Is it really fair to criticize the author for not having finished it yet?
Please look through the PDF. This isn't just not done. This is not ready to share with anyone publicly. There is no useful information in this. There are probably under 20 paragraphs of original text.
> Is it really fair to criticize the author for not having finished it yet?
No, but I'm criticizing the fact that it's posted[0]. Not that they're working on something.
I don't see the author here in this thread so my warning is to other readers. Just move on unless you're a book publisher looking for an author to pick up.
The only real criticism anyone could offer about this would be about the chapter structure, because that's all that exists. I would recommend they drop all the chapters that are a CS101 equivalent. There's no need to explain git or the OSI model or grep.
[0] edit, I want to clarify I mean just posted and dumped. If the author were here for questions or feedback I would feel differently. But with just this link as-is, there is no point in sharing.
Also data engineering focused, agreed completely.
The subtitle is "Mastering the Plumbing of Data Science"
> "If you are looking for AI algorithms and such data scientist things, this book is not for you."
> "How to use this document: First of all, this is not a training! This cookbook is a collection of skills that I value highly in my daily work as a data engineer."
> "You are going to find Five Types of Content in this book: Articles I wrote, links to my podcast episodes (video & audio), more then 200 links to helpful websites I like, data engineering interview questions and case studies"
> "This book is a work in progress! As you can see, this book is not finished. I’m constantly adding new stuff and doing videos for the topics. But obviously, because I do this as a hobby my time is limited. You can help making this book even better."
Although when I think of a cookbook, I'm mostly interested in some re-usable snippets of codes that can be used again and again. A good example would be Chris Albon's site https://chrisalbon.com
For eg: - a recipe for splitting comma separated values in a column into multiple rows or - a data cleaning recipe that removes all unwanted spaces/trailing newlines/punctuations
I've searched but never come across a compilation of such re-usable code snippets anywhere. Would be glad if anyone has any resources like this.
I am currently using one Python process with Prefect for DAGs (https://docs.prefect.io/guide/), custom API queries and Elasticsearch indexing code in a batch processing style and it seems to be going fine to ingest a 2TB dataset with the ability to generate rich errors, retry/resume etc.
Besides maybe job listings, why would a real "plumber" who came from python and js NEED to dive into the whole Apache ecosystem?
But really, let the imposter syndrome go.
I'm a third-generation computer engineer and the first one with a degree. I'm half the engineer my father is or my grandfather was. Whatever title you put on your resume is just you marketing the services you deliver. If you can do the work of a Data Engineer, you're a Data Engineer. If you're really self-conscious about it, get a gcloud/aws certification.
The main advantage of using the data engineer label is that it describes what you do now so that others will understand. It doesn't describe your academic past, only your professional present.
What's more, titles like data/ML scientist/engineer are still too new and ill-defined for anyone to worry too much about the exact work implied or the underlying credentials. 'Engineer' is just an attempt to imply that the role involves more involvement in production than exploration.
Business intelligence consultant/developer was a blanket term used either for 1) people that can model and translate business requirements into data platforms/components; 2) poorly targeted recruiting; 3) more rarely, to describe people that work both in 'frontend' (reporting and dashboarding using tools like PowerBI, who are "data analysts" in newspeak) and 'backend' (ETL developer).
"Data Engineer" today is, sadly, too often synonym with "solving solved problems using an unnecessarily complex approach and toolkit"; and then there are the cases where the volume or complexity of the data actually justifies the cost of using "big data" platforms. Not to sound harsh; the article discussed here illustrates my point using "simple" unix tools: https://news.ycombinator.com/item?id=14401399
I don't think I ever see a PDF link online and think, "yes, this is the ideal format for ingesting this information."
I liked case of studies compilation thumps up.
The problem comes about that we can have peaks of 20k reads and 20k writes per second with strict guarantees on response time and data consistency. All at the same time needing to be kept consistent across multiple datacenters in several regions.
Your typical application won't hold up in that environment or can become extremely difficult to maintain. And I've met people at conferences who would say my use case still isn't "big data" and is pennies compared to their data streams. It really does become a class of algorithms and solutions on its own, just dealing with your real-world ability to manage that much data.
Highly scalable distributed systems are the bread and butter of backend engineering in modern tech companies. Even specialists working on storage layer tech in particular are not called "data engineers," just backend engineers, although with more systems/infrastructure focus than others.
People with the title "data engineer" at my company pretty much do ETL pipelines on Hadoop. I've done some ETL pipelines on Hadoop; seems like a tool that should be in the portfolio of any SWE. But when you give a data engineer a problem, they are sure to suggest an ETL pipeline in Hadoop; I have even seen them build RPC interfaces out of INSERT statements. Whereas a SWE is looking at a broader suite of options, and probably defaulting to backend services / OLTP databases. Hence my confusion. Maybe our use of the term is anachronistic? Or maybe they just know a bunch of HiveQL or pipeline design patterns that are too complex to fit into the heads of people who can also write services?
> But when you give a data engineer a problem, they are sure to suggest an ETL pipeline in Hadoop
Maybe someone who sees Hadoop as their only hammer but IMO a "good" data engineer would go
"How much data? 300MB? And you need to run this how often? Once? You're sure just once? Like really sure? Okay here is a bash script."
Edit: Added clarification.
[1] https://www.amazon.com/Designing-Data-Intensive-Applications...
When we reviewed the design I mentioned a few points that were like: "Yeah I know the little requirements you gave counteracted this design, but if we do it this way it'll help us out (source in book)"
This book is really well written, and I've learned so much from it and I keep opening it up every day for further guidance.
After reading the book, googling for something like "mongodb v.s. cassandra" would start to feel as silly as googling for something like "javascript v.s. css" as you start to understand the fundamental differences between them.
No more need to hope the vague Medium post you found while trying to decide which DB to use would match your use case closely enough.
So your options are either senior software engineers who have done some data work (that's how I got to be a Data Engineer) or people who've been doing analytical data work (either in the traditional warehousing space or via science/insurance/finance type spaces) that are semi-technical but have no formal engineering background.
The former are people who went to college in the late 90s/early 2000s (like myself) when things were different. The latter need to hyperfocus on coming up to speed in engineering.
I reviewed this guide a couple months ago for my employer to consider as the basis of an internal bootcamp, and I'd note that it's perfect for the audiences I mentioned. Also, even for people with more up to date academic experience, note that the transactional database schemas that software normally deals with often look wildly different than analytical structures.
I have kept up to date on these technologies, I participated in undergrad research on distributed systems and my career has revolved around them. Many devs never really get a say in where their data goes, they might read a blog post or two about new systems, but it leaves a very light imprint. Its been rather spotty as to whether I had any say in where my data is stored throughout my career.
e.g. How might I go about optimizing a redshift query? Well, now that I have an idea about how data is laid out on disk, because redshift is a columnar store, if I try to optimize X query, here's how I imagine the index to be so that sequential reads would be faster.
I could find a reference on how to optimize redshift queries, but this book answers the WHY and not just the immediate how.
I've read so many books that were practical, yet became so much less useful over time. (e.g. reading a book about the specifics of the Angular API, whereas now I write mostly React.)
I keep returning back to this book for understanding a top-level view of the fundamentals of distributed systems, specifically data stores.
I hope you give the book a second look at some point.
I'm not a fan of the marketing gravy people dump on everything. The future will devolve to unhelpful titles like - "How to do Data Fking science"
Like how almost no recipe book would tell you that it's customary to serve cake on birthdays or that you should be careful of the high sugar, or various cake alternatives, you would just find a section for cakes and it assumes you're making the correct choice in searching for a cake recipe.
For example I own a cookbook on NLTK and flipping to a random page I got "How to remove insignificant words from a sentence"