Engineers Shouldn’t Write ETL
multithreaded.stitchfix.com
multithreaded.stitchfix.com
I sometimes gave my own suggestions on how to improve upon their ideas, but for the most part, I was happy to focus on implementing their ideas, in the most clean, elegant, robust and testable manner possible. I was happy to do the "plumbing" work of improving upon our tech stack and architecture, in order to make the entire system better functioning and easier to maintain.
According to the author, I'm supposed to resent the fact that I'm a "doer/plumber", and not a "thinker". In reality, it was the opposite. Do I really want to spend my entire day reading the Bloomberg manual and figuring out which tables/columns will give us the data we want, and the nuances of what this dataset does and does not cover? Sorry, I have zero interest in doing that.
I enjoy programming. I enjoy system design. I enjoy building stuff. I have zero interest in becoming an expert on how to interpret the Bloomberg symbology file. Besides, if I ever left the financial industry and joined a tech company, that knowledge will become completely useless.
Did I or anyone consider myself to be a "menial" plumber? I don't think so. I was getting paid hundreds of thousands of dollars, because the "thinkers" recognized the value that I brought to the table. They appreciated that I could quickly and robustly implement the ideas that they had, and keep the system running smoothly without hiccups. They recognized anyone can do a "good enough" job, but it's much much harder to find someone who can do a great job. And for my part, I was perfectly happy to be that guy.
If you're someone who wants to expand your breadth and take on more "thinker" responsibilities, more power to you. But just don't forget that there are people like me out there too. There's no shame in being an excellent "doer".
{
"Unit test" : "google mock",
"CSV" : "Apache Parquet",
"message" : "google protobuf",
"queue" : "Apache Kafka",
}
etcHis data science recommendations looked like a markov chain of various analysis algorithms.
One time I started digging into his recommendation trying to figure out why it was even on topic and he starts going on about how he's a "big picture" guy and not to bother him with implementation details. The thing was, his 'big picture' ETL was breaking our trading system every other week due to some inane dependency strictness that wasn't necessary.
There's nothing more 'big picture' than not fucking trading!
I guess because portfolio managers aren't specialized in engineering, we see a lot more of these fakers in finance than tech.
One of the reasons I love making people write down their idea (myself included) before talking about it is that writing forces a small initial execution step. Even a step this small can often filter the useless ideas away.
Wow, true system engineering! Your company sounds like a good place to do some professional work.
Sorry for my grammar above as I was typing on my mobile phone halfway on the bus.
I serve as a sounding board a lot for my oldest son and that works well, but it's not uncommon for such people to just be trying to meet their own need to process information and/or feed their ego, oblivious to how it impacts other people and not really welcoming of the feedback they really need for this to be constructive. A good sounding board doesn't just listen, they ask pertinent questions and make insightful comments that help move the thought process along.
Sometimes when I meet people like that, I'm able to direct the conversation to a more constructive back and forth of that sort. But some people just know they have this need to talk, they have a lot of baggage that makes them openly hostile to meaningful feedback and they crave validation. Anything other than praising their half-baked ideas is met with toxic reactions. In such cases, the best you may be able to do is basically make a few polite noises and then disengage as quickly as possible.
> they have a lot of baggage that makes them openly hostile to meaningful feedback and they crave validation. Anything other than praising their half-baked ideas is met with toxic reactions. In such cases, the best you may be able to do is basically make a few polite noises and then disengage as quickly as possible.
would serve pretty well as a fair description of normal people, ime.
For example, in grade 5 I got a C+ in fine arts. It devastated me to the point where I questioned my abilities and disengaged from schooling. Permanently. It's only been in the last few years that I've actually been able to apply myself to anything.
Now, I can see that was an unreasonable reaction. At the time though, I didn't even understand what was happening.
If I hadn't got a lot of help, I believe I would still be driven by my insecurities to this day and would be impossible to work with.
Edit: typo
My thought as well. I also worked in hedge funds for a long time, and I kept getting resistance from the self-proclaimed "thinkers" to do even the most basic project management things like keeping a shared list of bugs, using version control, etc. It became clear that they simply didn't know how to use these tools, while claiming to specialize in financial modelling.
This also turned out to be false, as moving on to other funds I discovered their way of seeing things was quite limited. Which I had a good suspicion of, but other people actually showed me how one could approach things.
Part of that reckoning was that to be good at building financial strategies -something virtually nobody will tell you about- you need to be fairly good with writing code. Not just your Frankenstein of VBA, Excel, and Matlab, but including a fairly deep understanding of algorithms as well as common DevOps tools.
I'm joking (sort of) but this is half the job description for leadership jobs, politicians, executives and such.
The differences are of quality, not type. Jobs' thinkering was good, for the most part, for example.
There are some forward-moving, solid "soft skill" analysts/data scientists that can make this happen. But by and large they shouldn't all be held to this standard. Maybe my standards/expectations have been soured too much and I'm too pessimistic, but as a whole they're just not cut out for this kind of stuff. Which is fine - being a "doer" is easy to begin with, and over time the more that you're able to automate as a data engineer, the more trivially easy ETL/everything else becomes.
I'd definitely not want to be stuck doing stage 2 forever, would prefer 3. I think you're saying that you enjoyed a job which was some 1 and some 2. I'm sure there's someone out there who wants to wire up pipelines with no engineering and no analysis all day but I'd imagine it's a rare breed.
Edit: I think the important distinction is team/company size. Doing a bit of everything as a 1 man team is challenging, if you have a team where devops/engineering/reports/tools have been chosen/built/standardized by specialist and you really are just wiring pipelines up, I think that would be tough. On the other hand being in a small team condemns you to always be doing the same fractions of work because there's noone to hand off to.
I think this is the key idea there. It's good that you found a situation where you're both appreciated, and compensated for it. It's too easy for engineers to be devalued as replaceable cogs in the pipeline of things that need to happen to bring in revenue.
Finance may have been ahead with adopting data science but it seems that many others reached parity or even surpassed in capabilities.
"How do I introduce an abbreviation in the text? The first time you use an abbreviation in the text, present both the spelled-out version and the short form." https://blog.apastyle.org/apastyle/abbreviations/
A team of one data scientist and one engineer, completely responsible for building a model, and seeing it through into production, meeting all applicable SLAs and performance metrics.
Or maybe it's two data scientists and one engineer, or one scientist and two engineers, whatever is required.
The point is to have a small team you can hold completely accountable for their output. They sink or swim together, so there is no debating whether the scientists or engineers get the credit or take the blame. They are assessed by the effectiveness of the end product they produce.
I was able to work with someone apt at machine learning while I focused on building out the UI and backend. We delivered a first release about 3 days after we started, giving ample time to seek feedback and let the users shape the direction.
No team needed or proof of concepts. Actual working data models, up to date tables + ETL code in production.
Background is DBA.
> We are not optimizing the organization for efficiency, we are optimizing for autonomy.
Efficiency is for production pipelines where the product is thoroughly defined and production costs eat deeply into profit margin. Most software organizations have massive margins - but only if they get to the right product. Organizing people for ownership and autonomy engages their creativity, but also ensures that the org can move forward even when one side or the other falls behind.
I think this is more the point than “engineers shouldn’t write ETL”: the engineering-related department consuming the ETL’s output should likely be the ones writing/maintaining it. Or, perhaps more generally: don’t delegate entirely to another team if the team that cares about the result is capable of meeting their own needs.
On the other hand, sometimes you can't get away from that because different orgs/humans generate trash data in idiosyncratic forms. Things will get much better once we pry all the human hands off of data and let engineers redesign all of them across the world. Not going to happen soon.
The person doing modeling or data analysis should ideally be dealing with the raw data, know how it was collected and understand what each field really means.
This is exactly the author's point. The Data Scientists are consuming the ETL's output, so they should learn how to write and maintain ETL since it isnt very hard or time consuming with modern tools.
That is its own soul-sucking experience. "Can't we just hire people that can learn Power BI better? Why are we still writing data tools for people that think they know Access but barely know Excel?"
https://www.stitchfix.com/careers?gh_jid=1252958&gh_jid=1252...
It's 2018. A lot has changed since 2016. The line between sw engineering & data engineering is much thinner.
When the landscape for tech in DE was oracle, mysql and Cognos, DE's didn't need to know about OOP or consensus algorithms. Because the landscape now includes hadoop, redshift, kafka, spark, airflow, notebooks, TiDB and lord knows what else, DE's need to have most of the skills of a software engineer to be successful.
Different parts of the org with different skillsets and cultures practicing empathy for each other by communicating interests in version-controlled code, allowing for guard-railed autonomy, which leads to business agility.
Yep. Sounds about right.
> Optimize for autonomy not efficiency
Optimizing for efficiency without considering the cost of work in progress (WIP) (irrelevant ETL models), rework (unscalable models), or unplanned work (unscalable models that make it to production) results in company silos (data engineering, infrastructure engineering) cheering local maxima while covering their ass in the face of a business that's suffering from a long lead time. Two teams with two backlogs will accomplish work exponentially faster compared to three teams with three backlogs.
It boggles my mind how books like The Phoenix Project are not required reading.
As fun as it is to build with and learn new technologies, it’s a bad idea to build data pipelines unless you have a lot of resources and good leadership that can make peace between all the different people who touch the data.
Unfortunately in the world of sensors and equipment there aren’t many solutions, so I started a company (at https://sentenai.com ) to save others from my years of struggle. It turns out it’s even harder to build a general time series data pipeline solution, but we’re making progress.
The startup I'm employed at needs some data analysis, but it is not big data, simply a way to unify analytics into a queryable database. I'm not looking forward to writing any ETL code, and was hoping someone here had a tool to help.
Internally, we're using a tool called Meltano which is aimed at solving just your problem. Most of our data warehouse is coming from external business ops tools (Salesforce, Zuora, Zendesk, Marketo, etc.) and we're using dbt for transformations w/ Looker as the BI layer. Definitely check our primary analytics repo [0] as all of our code is out in the open. Feel free to ping me if you have more questions - tmurphy at gitlab.
I would say that you should pilot with a few ETL vendors. We currently use Fivetran, they're fine but we've had enough burps that I cannot cold recommend them over other vendors. I cannot for the life of me remember the details, but I think we went with them over Stitch for pricing reasons.
You're totally right about all of the flavors DB2. Our support team would be the ones to figure out for sure whether or not we can work with your setup, and you reach them at support@stitchdata.com
This does not strike me as a great idea.
If any of you older/more experienced engineers and scientist have advice or wisdom for me, I would very much appreciate it.
Don't worry too much about "falling behind". There will always be time to learn more math or a new framework. Worry more about finding that first job, any job, then you can branch out once inside the industry. Networking beats recruiters beats sending a resume, so try to find a friend who already works where you want to be.
I would love to do my own startup. I have a few ideas floating around. But I feel like I lack the discipline to sit down every day and force myself to work on them without external deadlines/pressure.
In terms of jumping into the tech industry: I understand the advice about looking for any job when starting out. It just seems that even a lot of the entry level jobs are very specialized.
2. Look for work in a different part of the country... or maybe just the right organization that’s willing to take a chance on you. We’re in Austin and we’ve hired smart, hardworking kids who’ve never touched the languages we use and get them contributing meaningfully in <2 months.
3) I used to work for a startup where our CTO would not hire a data scientist unless they could write production code (in backbone and rails which I didn’t know at the time) and after I started, I spend 4-5 months just learning to be a full stack dev- I think that was so useful for my career as a data scientist. It meant that data scientists at this company could put whatever models they were running into the product- it drastically simplified org structure—- much more autonomy and fewer project management dependencies.
Sure not a lot of data scientists wanted to also be or make them selves into full stack developed, but you’d end up with the really gritty ones who and they’d end up being more loyal and much more on the same page as the rest of the engineering team it was way better for the whole org.
We’re hiring interns + junior full time people by the way
It's ridiculous to me how hard it is for him to find a new job at 60. He financially doesn't have to, but he wants to train younger guys on how to deal with all the weirdness one encounters in ETL Jobs.
Is this the common perception, because it really doesn't line up with my experience?
Writing/updating a report is easiest part of my job it's the data that goes into building it that is hard translating the "simple metric I've developed" and getting it to run in a robust automated and sane fashion is the difficult part.
The complexities in my org are two fold.
Firstly the infrastructure people don't get data - at all. They speak PLC's and HMI's to them it's all OPC and magic A2A messaging takes care of everything. All data is time series to them and it all goes into an historian (which is basically a giant ring buffer i.e it gets flushed periodically) anything beyond that is past their level of expertise.
The data needs to be batched together the time series information has to be processed into "event frames" - this data was all part of this sequence of conveyor belt movements for example. Then you need to link it to related events etc and archive it in some kind of sane fashion so that in six months time if there is a product defect or something like that you can trace the entire series of event frames for that particular production batch.
Secondly the people the article calls "data scientists" (in my org these are Engineers - real ones of the Chem and Mech variety) don't know anything about databases or handling data they prototype their metrics in Matlab, Fortran, Excel and the like.
You really need someone to translate their code into something sane that can be automated. Engineers are not taught to code at all. I know I studied engineering at university Fortran is the lingua franca. Code is just a way of representing mathematics. Asking these people to do all the data processing pipeline is just not going to happen. It's not their job. They write the simulations and models they have the domain knowledge thats whats important for them to be worrying about.
Reports that go externally are done by certified people. (Laboratory technicians for product specifications and finance analysts for stock market stuff).
I think the business hires data scientist to be informed. Not to make business decisions on their behalf.
> Data scientists love working on problems that are vertically aligned with the business and make a big impact on the success of projects/organization through their efforts. They set out to optimize a certain thing or process or create something from scratch. These are point-oriented problems and their solutions tend to be as well. They usually involve a heavy mix of business logic, reimagining of how things are done, and a healthy dose of creativity
Again, I'm confused? That sounds like the data scientists should have majored in business then. If data scientists start doing that, what will all the other business folk do then?
Data scientists should just build out reports that provide valuable insights and potential patterns that can help make business decisions. The difference with prior reports engineer or data analysts or wtv, is that a data scientist is assumed to be able to generate statistical analysis or/and pattern analysis over the data. While prior, a data analyst only needed to perform basic versions of that which did not go beyond what SQL could do.
The data engineer should enable the data scientist to perform this analysis by both working with the software engineers to acquire it safely, securely, reliably and at scale. And working witj the data scientist in order to apply his statistical analysis efficiently and at scale to a possibly very large data set. Finally, he might need to work with both software engineer and data scientist to setup real time or close to real time versions of the analysis.
All result from the analysis should be presented (aka reported) to the business. The data scientist can suggest interpretations or ideas to address findings, but it's the business role to make tactical and strategic decisions about business processes and products.
And if you're doing ML as part of a process, then you need a ML scientists. Say you need to build out voice recognition, or the likes. Basically comp sci or math majors with ML masters or PHDs.
https://s3.amazonaws.com/xplenty-assets/infographics/raw_dat...
This particular line really rubs me the wrong way.
Do the best you can... You won't be the best in the world, but you can still have a positive impact.
What this article is really saying is that replicating your data from source apps shouldn't be manually coded. The harder part still needs someone to write code so business users don't need to.