HNHacker News
TopNewBestAskShowJobs

adrianbr

11 karma · joined October 11, 2023

submissionscomments
adrianbr··on Show HN: Melchi – Open-Source Snowflake to DuckDB Replication with CDC Support
The whole idea is pretty nuts! I can imagine it being used as the dev env for teams that use SQLMesh and can thus port the sql. Might be worth investigating with them
adrianbr··on Koheesio: Nike's Python-based framework to build advanced data-pipelines
That's really cool, did you already saw the dlt library? That one's done for very easy to use EL in python. It's similarly modular and built by senior data engineers for the data team, and the sources are generators which you could probably use too.

How is koheesio different to dlt? Where could they complement each other?

adrianbr··on Show HN: Hamilton's UI – observability, lineage, and catalog for data pipelines
congrats on the hard work and this launch!
adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
ahh good old manual fine tuning and maintenance. We are adding data contracts for things like event ingeston where schema needs to be strict or cases where you know ahead of time what to expect.

Our experience comes from startups that usually do not have time to track down the knowledge and rather go out and find/make their own. Here you definitely want evolution with alerts before curation - so load to raw, and curate from there. Picking out data out of something without a schema is called "schema on read" and you can read about its shortcomings. So this is both robust and practical.

For the fine tuning, as I mentioned, data contracts are a PR review and some tweaks away. They will be highly configurable between strict, rule based evolution, or free evolution. Definitely use alerts for curation of evolution events!

adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
Thank you! That's the example we looked at for our dlt-airflow integration :) the dlt dag becomes an airflow dag.
adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
Duckdb is analytical and gained popularity with the analytics crowd. it has multiple features that make it play well with use cases in that ecosystem such as aggregation speed, parquet support, etc
adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
for now :) Thanks for pointing it out - and it looks like we should add an aws lambda guide too :)

If you want to deploy to lambda, try asking in the slack community, some folks there do it.

Or if you wanna try yourself, here is a similar guide that highlights some concerns from deploying on gcp cloud functions https://dlthub.com/docs/walkthroughs/deploy-a-pipeline/deplo...

adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
Thank you for the heads up!

it is unfortunate, and with 3 letter acronyms this will happen.

An easy way to remember is that we are the one you can pip install and plays well in the ecosystem.

Databricks has interesting choice in marketing names, ngl - DLT named after the competing dbt - Renaming standards like raw/staging/prod to Bronze, Silver, Gold

adrianbr··on Show HN: OpenAPI DevTools – Chrome extension that generates an API spec
This is amazing! to figure out the website apis has always been a huge pita. With our dlt library project we can turn the openapi spec into pipelines and have the data pushed somewhere https://www.loom.com/share/2806b873ba1c4e0ea382eb3b4fbaf808?...
adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
There are multiple ways to run together - we will show a few in a demo coming out soon.

We also consider a tighter integration like with Airflow described here as a possible next step https://dlthub.com/docs/walkthroughs/deploy-a-pipeline/deplo...

We will investigate the interest incrementally as to not build any plugins that don't end up used.

adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
you can request a source or a feature by opening an issue on sources/dlt repo https://github.com/dlt-hub
adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
Since dlt generates a schema, and tracks evolution etc, contains lineage, and follows data vault standard it can easily provide metdata or lineage info to the other tools.

At the same time, dlt is a pipeline building tool first - so if people want to read metadata from somewhere and store it elsewhere, they can.

If you mean to take metadata like we integrate with arrow - that remains to be seen if the community might want this or find it useful, we will not develop plugins for collecting cobwebs, but if there are interested users we will add it to our backlog.

adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
Thank you for the feedback! I can see now how it could be confusing.

The reason we used chatgpt is because it's an easy starting point - why read through examples when you can get the one you want in seconds?

Because dlt is a library, it's closer to how language works and gpt can just use it - from our experiments, we cannot say the same about frameworks.

adrianbr··on Show HN: Dlt – Python library to automate the creation of datasets
dlt is a python library that you can probably plug into the OpenRefine java application to enable moving the data somewhere easily and into different formats, making OpenRefine more useful in a connected environment.

I would not say they are similar - rather OpenRefine is made for visual data cleaning, while dlt is made for automation of data movement with structuring and typing to enable crossing different format standards with ease.

Together you should have a good combination of automation and manual tweaking option if needed.