HNHacker News
TopNewBestAskShowJobs

albertstanley

5 karma · joined December 14, 2022

submissionscomments
albertstanley··on Launch HN: Serra (YC S23) – Open-core, Python-based dbt alternative
Sure, I’ll clarify some of the terms used in that one-liner in case it’s helpful for anyone else as well.

ETL is the process of extracting transforming and loading data from a source to a destination in a data pipeline. Spark, an engine for large scale data processing, allows us to write code that can work with large amounts of data. dbt is a tool you can use to break up your SQL scripts into smaller “models” - other SQL scripts that can be reused and tested.

We described us as an end to end because we also have extractors and loaders, whereas dbt focuses on the T ( transformation step of ETL ). Each of our steps involved in extraction, transformation and loading correspond to a specific Python object defined in our Python framework. I have also updated the README in our repo to hopefully better explain how the config file links to user defined readers, writers, and transformers.

albertstanley··on Launch HN: Serra (YC S23) – Open-core, Python-based dbt alternative
Sure, our approach is to define Python classes to handle reusable steps for reading, transforming or loading data. For example, we have a MapTransformer, CastColumnsTransformer, GeoDistanceTransformer.

Each class specifies some configuration needed for the "step" and can then be used in the config file to construct a full ETL job. You can write unit tests for custom transformers you create as we have shown in the tests/ directory.

I have also updated the README in our repo to hopefully provide a better explanation of how our config file connects to specific Python objects.

albertstanley··on Launch HN: Serra (YC S23) – Open-core, Python-based dbt alternative
This is a completely valid point, we'll be changing the readers to directly read into Spark. Thank for the comment!