Shopify's Data Science and Engineering Foundations (2020)
shopify.engineering
shopify.engineering
I've never understood why data science teams are typically so far removed from "normal" engineering teams. Maybe it's the DevOps kool-aide speaking, but in my opinion, teams should be more horizontal than vertical!
Pros of same-team: fewer ideas "lost in translation" between data scientists and data engineers, better understanding of which datasets/flows are top priority, can sometimes share some stack components and help datascientists improve their code, better chances of getting data scientists to contribute their own batch jobs (there's just more trust as opposed to dealing with some "engineering" team that is less connected to you)
Cons of same team: data engineers may not be as in-the-loop on what's happening with production datasets, may not be as tightly integrated with a devops team, may get overly caught up in "business logic" as opposed to "plumbing".
When it's two independent teams, you tend to get a more research focused Data Science organization and a team of engineers more focused on plumbing.
Which option is better will depend on the organization goals. If you think that you have a straight research problem than a dedicated research team is useful. If you want to ship product than one team is better.
> We test every situation that we can think of: errors, edge cases, and so on.
Pretty bold claim. Either their fantasy can't think of many weird data issues, or their data is somehow guaranteed to be always validated, or they came up with a magic secret sauce for data unit testing.Could anybody explain what they mean by their version of unit testing?
Does it mean to test that a change in code will still successfully process the same static data sample as was tested previously?
If not using static data, what does it test when the input data changes? Does it validate the input data?
Doesn't say anything about how many situations they think of.
It's surprising to see Shopify sticking with the Data Warehouse model in 2020… I expected something a little more cutting edge.
Data Warehouses are fine if you are working with a small number of data sources, but at some point they start to slow down development and new analytics, and make it cumbersome for intraday reporting.
This is why I'm seeing Data Warehouses being replaced by Data Lakehouses at my company and others in the industry. The Lakehouse enables faster development and near real-time analytics, on lower cost storage and works with structured and unstructured data. Similar team practices are still applicable, but the underlying data structures and governance is different.
Wish I could write a script to keep retrieving my Amazon order history with one click.
Anyone have any solutions you've tried?
It’s a great approach for your data lake and data warehousing needs.
If you want quality, you need structure and review. Accessible data is helpful and needed to develop some of the mature processes, but for most day to day analysis/reporting, no one wants to create their own data model from scratch.
Lots of FAANG doesn't apply to any other companies so it may just be a case of having a wholly unique use case. Though I'm surprised there isn't something already in place at this point (of course having very little knowledge of the case). For the dims/facts/marts, they tend to be business use case focused and not source/data which can reduce the targets down significantly since business use cases tend to repeat (or rhyme).