411 karma · joined April 8, 2018
To that end, have you happened to notice some of nbdev’s design choices bleed into adjacent tooling? Which pieces do you hope to endure and take new life in future projects?
What’s the biggest barrier keeping larger audiences from developing in this “literate” manner?
I’ve been a follower of the design philosophy of nbdev for a long time, though admittedly never got to use it. Perhaps now is the time! Cheers
praise be the algorithmic DJs that brought us Ethiopian Jazz in the 2010s. While there's certainly valid criticism of the payment model of music streaming platforms, I love how many "under-loved in their time" artists are now in the spotlight and appreciated as a result such as Astatke.
coincidentally, my PR to the dbt viewpoint was closed by the docs team as "closed, won't do" [1]
I really like the convention of data plane (where you describe how the data should be transformed) and the control plane (i.e. the configuration of the DAG, do this before this). In this paradigm, I believe that the control plane should be as simple as possible, and even perhaps limited in what can be done with the goal of pushing the user to take data transformation as tantamount. Maybe this is why I fell in love with dbt in the first place is because it does exactly this.
"spicy" take: allowing users to write imperative code (e.g. using loops) that dynamically generates DAGs are never a good idea. I say this as someone who personally used to pester framework PMs for this exact feature before. While things like task groups (formerly subDAGs) [2] appear initially to be right answer, I always ended up regretting them. They're a scheduling/orchestration solution to a data transformation problem
Can y'all speak to how Hamilton views the data and control plane, and how it's design philosophy encourages users to use the right tool for the job?
p.s. thanks for humoring my pedantry and merging this! [3]
[1]: https://github.com/dbt-labs/docs.getdbt.com/pull/2390 [2]: http://apache-airflow-docs.s3-website.eu-central-1.amazonaws... [3]: https://github.com/DAGWorks-Inc/hamilton/pull/105
so often w/ pandas I’d: 1. “yeet” the csv into a dataframe 2. use dataframe methods to massage the data to a “clean” state 3. push as much of the df methods into pd.read_csv() parameter options
it’s be great to iterate more quickly on the above loop. better yet — what if it would could auto-generate a letter to send to the folks from whom you got this data on how they could better output to csv to make ingestion simpler and easier for downstream users…. but maybe that letter would just be "don't use CSV!"
related to flat data formats, it obviously makes sense to start with CSV, but what about the future? If this tool became ubiquitous, how might a SWE or data professional's job change? What opportunities be created? As in:
1. CSV is ubiquitous but has no singularly well-adopted standard. 2. software and data engineers struggle with CSVs as a result of #1. 3. tool is created to reduce pain and friction. 4. profit? a new market? a new standard?
Last, but most personally interestingly, how much do you know about the Apache Arrow ecosystem and how it's mission might overlap with YoBulk's
> coined the term Online analytical processing (OLAP) and wrote the "twelve laws of online analytical processing". Controversy erupted, however, after it was discovered that this paper had been sponsored by Arbor Software (subsequently Hyperion, now acquired by Oracle), a conflict of interest that had not been disclosed, and Computerworld withdrew the paper.
fyi the hyperlink "found here" to setup_new_database.template is broken.