Great questions!
> "spicy" take: allowing users to write imperative code (e.g. using loops) that dynamically generates DAGs are never a good idea.
Can you give some examples of when it was a bad idea? Otherwise to clarify, with Hamilton, there is no dynamism at runtime. When the DAG is generated, it's generated and it's fixed. The operator we have for doing this `@parameterize` requires everything to be known at DAG construction time. It's really just short hand for manually writing out all the functions. So I don't think it's quite the same story - it is more a "power user feature" - but when used, it makes code DRY-er, at the cost of some code readability.
> While things like task groups (formerly subDAGs) [2] appear initially to be right answer, I always ended up regretting them. They're a scheduling/orchestration solution to a data transformation problem
Yep. Hamilton has a concept of `subdag` too. It's really short hand for "chaining" Hamilton drivers. We take the latter approach (I believe) since it is there to help you more easily reuse parts of your DAG with different parameterizations. Since Hamilton isn't concerned with materialization boundaries we don't have to make a decision here so how it impacts scheduling/orchestration can be punted to a later time :)
> Can y'all speak to how Hamilton views the data and control plane,
Hamilton is just a library. There is no DB that needs to be run to use Hamilton. The only state required is your code. So at the simplest micro-level, Hamilton sits within a task, e.g. creating features, and replaces the python script you'd run there. So at this level, I'd argue data plane vs control plane doesn't really apply, unless you view code as the control plane and where it runs as the data plane... At the macro-level, e.g. a model pipeline pulling data, transforming it, fitting a model, etc., you can logically describe the dataflow with Hamilton, without breaking it up into computational tasks a priori. I'd say Hamilton here tries to be agnostic and provide the hooks you need to help you coordinate your control plane and facilitate operation on your data plane. Note: this is where we see DAGWorks coming in and helping provide more functionality for. E.g. with Hamilton you don't need to decide whether everything runs in a single task say on airflow, or multiple. It's up to you to make that decision. The beauty of which, is that conceptually, changing what is in a task, is really just boilerplate given all the information you have already encoded into your Hamilton DAG.
> and how it's design philosophy encourages users to use the right tool for the job?
With Hamilton, we believe python UDFs are the ultimate user interface. By using Hamilton we force you to chunk logic, integrations, into functions. We also provide ways to "decorate" function logic which gives the ability to inject logic around the running of said functions. So we're really quite agnostic to the tool, but want to provide the hooks to be able to easily and cleanly add, adjust, remove them. For example, to switch between Ray and Dask, our philosophy is that ideally you can write code that is agnostic to knowing about the implementation. Then at runtime add those concerns in. As another example, the ability to switch/change say observability vendors, should not force a large refactor on your code base. We have an extensible `@check_output` decorator that should constrain how much you "leak" from the underlying tools. In short: (1) write functions that don't leak implementation details, they should just instead try to limit to just expressing logic; (2) the Hamilton framework should have the hooks required for you to plug in "tool" concerns. Does that make sense? Happy to elaborate more.