Koheesio: Nike's Python-based framework to build advanced data-pipelines
github.com
github.com
The idea that you'll just build a tool that makes hiring 10x as many inexperienced devs work is dubious. Just one more new DSL bro. Certainly we have cracked the code that no one else has.
The problem with these types of orgs/tools is that by its nature your DSL constrains the juniors/inexperienced devs to what is currently possible. There is not a lot of learning unless you rotate them through the tools team periodically, which no one does. It's also awful for the devs who are building in experience in something they can't use anywhere else.
I was in one shop where the "tools team" guys epiphany was he would meta-recruit by poaching the best data engineers out of the ETL team, lol. Very explicit "good team / bad team" vibes.
I find this type of thing scary as an outsider looking in. How a company so large has such immature engineering continues to astonish me.
It's management that doesn't want to risk their positions by doing the very difficult business of either starting over or properly simplifying their stack. It's not easy, it's not quick, but if they can't even answer that basic question then they need to do the work.
It's not usually an issue of immaturity, it's just really hard. To make things worse, often people don't really want to do the work because literally any other data engineering job would probably be more enjoyable.
Simplifying the tech stack would probably require simplifying their business operations, which probably means less revenue.
Starting over is often literally not possible because there are so many interconnected systems that aren't all necessarily owned by the company trying to make the decision...
I don't understand this perspective. Simplifying the tech stack might mean taking multiple services in multiple languages, and deprecating some in favor of migrating that functionality to the most maintainable codebase. This shouldn't mean "simplifying their business operations", or affecting their business operations in any way.
But I would imagine there are a lot of pieces of apparent cruft hanging around that is actually there because if you remove it things break.
Maybe a large retailer that you rely on requires an integration with an old version of SAP, and then a logistics partner only provides files over FTP, and you need to use OCR to retrieve any data from the files they're sending.
Management can't just mandate that you will 'simplify the tech stack'. Even refactoring smaller parts of the tech stack is often a pretty massive job.
> Management can't just mandate that you will 'simplify the tech stack'.
In my experience, it's the devs who want to do reactors for love of the craft and satisfaction of a job "well" done, while management seeks to sweep as much technical "debt" as possible under the rug, and they'll burn through as many Romanians' nights' sleeps as it takes to reach their numbers.
That is perhaps not a great example. My brother is a business analyst at Nike (has been for 15 years or more). I just asked him how hard it would be to answer that question and he said it would be pretty easy. Granted, this is the kind of data he works with routinely, so it may be more difficult for other teams that do not.
I worked at a shop that did this and the trade-off is TTM, as your 2 person tools team is constantly needing to unblock ETL team with new features as they encounter new requirements in the wild.
If your ETL team is 20+ people and the tools team doesn't have a head start, tools team will quickly fall behind an insurmountable backlog as your ETL team spins its wheels. But you might save some money if you choose the right KPI..
I walked into 20 years of adhoc code that had zero data lineage, recoverability, or scalability that was breaking daily. There were contractors with over a decade of tenure whose job it was to troubleshoot and fix their own brittle processes (and make new ones) daily.
I got laid off (747Max plus pandemic) as I was rolling it out and they went back to the old way.
Subsequently, a new startup (Pantomath) emerged with former GE engineers (and other former colleagues of mine) from my former department to address that problem domain.
Based on my experience trying to socialize this type of solution, sales are going to be a bitch.
From the readme they describe it as "robust" and having a high level of type safety, so I'm guessing they're just leaning towards the "learn about bugs up before they hit production" end of the spectrum than you.
Then again I don't do any of this data engineering stuff so maybe it doesn't matter too much if it doesn't work reliably?
I think this pretty much sums it up, "a structured way".
It's looks to be a thin wrapper around spark to provide a consistent way to structure ETL jobs. They've implemented a mini dsl defining jobs as a datastructure on top of Spark.
I've seen several companies build stuff similar to this internally, defining jobs as a data structure. It all amounts to each company having their own internal conventions, their own view of what is easier for their devs and creating a framework for it. Nike have just decided to make theirs public.
You can do all this simply with simple spark scripts. Personally I'd use simple spark scripts. Big companies with lots of people love making these tools as companies love conventions, their conventions, their style guides, deal with staff churn/on-boarding frequently so believe these kind of things make that easier.
Probably makes sense in Nike as a way of organizing their ETL jobs but that looks to be all it brings. A way to structure/define your simple spark jobs the way Nike devs believe it should be done.
seems to be Spark-only AFAICT?
https://engineering.nike.com/koheesio/latest/tutorials/onboa...
I’m not against proprietary software, but your website still advertises this product as an open source ELT.
Yes, it is, because nobody wants to run multiple orchestrators, and the "What sets Koheesio apart from other libraries?" section does little to help users decide why they should pick yours.
Workflow orchestration is a mature category, as evidenced by the length of this list: https://github.com/meirwah/awesome-workflow-engines
I would expect someone who's seriously writing a new orchestrator in 2024 to cite the alternatives, their shortcomings, and how you intend to address them. Bonus points if you make a neat little table.
The fact that you're leading with Python does not inspire confidence. Pretty much all workflow orchestrators use Python for their glue, and that's hardly the interesting part.
What were they using at Nike before this?
For every n-th solution in the market, n-1 existing ones could have been used, but they weren't for (many times) good reasons.
And looking at their Makefile and pyproject.toml I can see that they knew what they wanted.
Like it or hate it, it’s how social media works. Everyone posts other people’s things, and then more people get together to rip it to shreds.
If the author can clearly show the value proposition of their library, it will get more adoption, and the community and the author will get value.
If the author realizes they coded something that is already a well-solved problem, or a poorly-constructed alternative, the author and Nike could gain by throwing this away and going with a better alternative.
Personally, I wasted too many hours of my life creating a solution that did not solve any problems. Had I done some thinking and/or market research ahead of time, it would have saved me ton's of time to work on more worthwhile endeavors.
Good feedback is worth its weight in gold, even if it hurts a bit to hear it.
Straw man argument.
>I would expect someone who's seriously writing a new orchestrator in 2024 to cite the alternatives, their shortcomings, and how you intend to address them. Bonus points if you make a neat little table.
Did you even try to read the docs before you launched this critical diatribe?
From the docs (https://engineering.nike.com/koheesio/latest/tutorials/onboa...):
Advantages of Koheesio
Using Koheesio instead of raw Spark has several advantages:
Modularity: Each step in the pipeline (reading, transformation, writing) is encapsulated in its own class, making the code easier to understand and maintain.
Reusability: Steps can be reused across different tasks, reducing code duplication.
Testability: Each step can be tested independently, making it easier to write unit tests.
Flexibility: The behavior of a task can be customized using a Context class.
Consistency: Koheesio enforces a consistent structure for data processing tasks, making it easier for new developers to understand the codebase.
Error Handling: Koheesio provides a consistent way to handle errors and exceptions in data processing tasks.
Logging: Koheesio provides a consistent way to log information and errors in data processing tasks.
In contrast, using the plain PySpark API for transformations can lead to more verbose and less structured code, which can be harder to understand, maintain, and test. It also doesn't provide the same level of error handling, logging, and flexibility as the Koheesio Transform class.
It took me less than 15 seconds to find the solution to the problem you propose. How long did it take you to formulate your critique? Do you perhaps just have a prejudice against Nike (corporate haze), or is it an investment in a 'competing orchestrator' that is clouding your judgement?Also, the fact that they say that the alternative is "raw Spark" tells me either that they're confused, or not very good at explaining. Spark is used to execute tasks in a pipeline, not to orchestrate it.
Flyte went OSS what, 4 years ago? I'm not super familiar with it, but a) could have been that it was too unpolished at the time or b) requiring K8s to be a non-starter for some teams/ orgs. Same for Kubeflow.
We also don't know for how long Koheesio existed within Nike.
In short, there's a lot we don't know and there's a good chance the internal reasoning for investing in this made sense under certain circumstances in the past.
Doesn’t mean the project originated two weeks ago.
Or could it be an evidence that the existing tools all have their flaws and any reasonably sized organization will hit these flaws pretty early, so many orgs come to the conclusion that the best approach is to just roll their own that fits their case relatively well?
If so, I think you would recognize how unreasonable the presumption of your comment is. It certainly strikes me as such.
Most if not all the big OSS frameworks and commercial offerings originated in one big corp and subsequently moved into the Apache direction (Airflow) or spun out into their own companies.
As far as I can tell from the docs, this is not an orchestrator at all. It is a Python-wrapper for certain runtimes like PySpark. I don't see anything in the docs that mentions DAGs, dependency definitions, scheduling, or deployment.
How is koheesio different to dlt? Where could they complement each other?