11 karma · joined October 11, 2023
How is koheesio different to dlt? Where could they complement each other?
Our experience comes from startups that usually do not have time to track down the knowledge and rather go out and find/make their own. Here you definitely want evolution with alerts before curation - so load to raw, and curate from there. Picking out data out of something without a schema is called "schema on read" and you can read about its shortcomings. So this is both robust and practical.
For the fine tuning, as I mentioned, data contracts are a PR review and some tweaks away. They will be highly configurable between strict, rule based evolution, or free evolution. Definitely use alerts for curation of evolution events!
If you want to deploy to lambda, try asking in the slack community, some folks there do it.
Or if you wanna try yourself, here is a similar guide that highlights some concerns from deploying on gcp cloud functions https://dlthub.com/docs/walkthroughs/deploy-a-pipeline/deplo...
it is unfortunate, and with 3 letter acronyms this will happen.
An easy way to remember is that we are the one you can pip install and plays well in the ecosystem.
Databricks has interesting choice in marketing names, ngl - DLT named after the competing dbt - Renaming standards like raw/staging/prod to Bronze, Silver, Gold
We also consider a tighter integration like with Airflow described here as a possible next step https://dlthub.com/docs/walkthroughs/deploy-a-pipeline/deplo...
We will investigate the interest incrementally as to not build any plugins that don't end up used.
At the same time, dlt is a pipeline building tool first - so if people want to read metadata from somewhere and store it elsewhere, they can.
If you mean to take metadata like we integrate with arrow - that remains to be seen if the community might want this or find it useful, we will not develop plugins for collecting cobwebs, but if there are interested users we will add it to our backlog.
The reason we used chatgpt is because it's an easy starting point - why read through examples when you can get the one you want in seconds?
Because dlt is a library, it's closer to how language works and gpt can just use it - from our experiments, we cannot say the same about frameworks.
I would not say they are similar - rather OpenRefine is made for visual data cleaning, while dlt is made for automation of data movement with structuring and typing to enable crossing different format standards with ease.
Together you should have a good combination of automation and manual tweaking option if needed.