HNHacker News
TopNewBestAskShowJobs

MatthausK

46 karma · joined February 9, 2012

submissionscomments
MatthausK··on A Preview of DuckDB v2.0
The CEO/Co-Founder of dltHub/dlt here.

For our community DuckDB is the default data warehouse for local development environment. Last month +90,000 users used dlt (and their AI code editor) to load data into DuckDB.

Because of our proximity to the DuckDB community we are seeing enterprise DuckDB usage first hand. People imo sleep on the data volumes DuckDB can handle. We see Fortune 100 companies use dlt and DuckDB in production on their Lakehouses in hybrid cloud deployments. I can eg mention Stellantis (Chrysler, Jeep, Peugeot etc) because they talk about it publicly.

MatthausK··on Show HN: I built an open-source data copy tool called ingestr
one of the dltHub founders here - we aim to address this in the coming weeks
MatthausK··on Show HN: I built an open-source data copy tool called ingestr
one of the dltHub founders here - we aim to address this in the coming weeks
MatthausK··on Show HN: Dlt – Python library to automate the creation of datasets
Pulling from and into production databases is one of the early favourites from our dlt user base. Some reasons explained here in this MongoDB example (https://dlthub.com/docs/blog/MongoDB-dlt-Holistics)
MatthausK··on Show HN: Dlt – Python library to automate the creation of datasets
We hear a lot about the dlt & AWS Lambda. We have currently one user working on the use case (see our Slack https://dlthub-community.slack.com/archives/C04DQA7JJN6/p169...)
MatthausK··on Show HN: Dlt – Python library to automate the creation of datasets
Thanks for your vote of confidence & support Max!
MatthausK··on Show HN: Dlt – Python library to automate the creation of datasets
We took at least one immediate practical good piece of advice out of this which is that we should release a conda package and make sure that dlt works in it.
MatthausK··on Show HN: Dlt – Python library to automate the creation of datasets
1) Yes. We support all the databases and buckets as data sources as well. Some examples: - get data from any sql database: https://dlthub.com/docs/dlt-ecosystem/verified-sources/sql_d... or https://dlthub.com/docs/getting-started#load-data-from-a-var... - do it super quickly with pyarrow: https://dlthub.com/docs/examples/connector_x_arrow/ - get data from any storage bucket:https://github.com/dlt-hub/verified-sources/tree/master/sour... 2) Strictly technical answer: on the code level sources and destinations are different Python objects so the answer is no:) but you as a user rarely deal with them directly when coding
MatthausK··on Show HN: Dlt – Python library to automate the creation of datasets
You can use pydantic models to define schemas, validate data (we also load instances of the models natively): https://dlthub.com/docs/general-usage/resource#define-a-sche...

We have a PR (https://github.com/dlt-hub/dlt/pull/594) that is about to merge that makes the above highly configurable, between evolution and hard stopping: - you will be able to totally freeze schema and reject bad rows - or accept the data for existing columns but not new columns - or accept some fields based on rules'