> acquiring and moving data around,
yep with Hamilton we provide the ability to cleanly separate bits of logic that's required to change and update this. For example you'd write "data loader functions/modules" that are implementations for say reading from a DB, or a flat file, some vendor. If they output a standardized data structure, then the rest of your workflow would not be coupled to the implementation, but the common structure which Hamilton forces you to define. That way you can be pretty surgical with changes and understanding impacts.
Regarding assessing impacts, Hamilton provides the ability to "visualize" and query for lineage as defined by your Hamilton functions. We think that with Hamilton we can make the "hey what does this impact?" question really easy to answer, so that when you do need to make changes you'll have more confidence in doing them.
> managing the data over the lifetime of a model’s use,
Hamilton isn't opinionated about where data is stored. But given that if you define the flow of computation with Hamilton and use a version control system like git to version it, then all you need to then additionally track is what configuration your Hamilton code was run with, and associate those two with the produced materialized data/artifact (i.e. git SHA + config + materialized artifact), you have a good base with which to answer and ask queries of what data was used when and where. Rather than bringing in 3rd party systems to help here, we think there's a lot you can leverage with Hamilton to help here.
For example, we have users looking at Hamilton to help answering governance concerns with models produced.
> and serving adjacent needs (e.g. post-deployment analytics).
If it's offline, then you can model and run that with Hamilton. The idea is to help provide integrations with whatever MLOps system here to make it easy to swap out.
For online, e.g. a web-service, you could model the dataflow with Hamilton, and the build your own custom "compilation" to take Hamilton and project it onto a topology. During the projection, you could insert whatever monitoring concerns you'd want. So just to say, this part isn't straightforward right now, but there is a path to addressing it.
While Hamilton/DAGWorks is mainly for expressing the pipeline/abstracting away the infrastructure, DAGWorks can help make the model lifecycle easy as well:
- Hamilton pipelines can be run anywhere, including in an online setting
- Breaking into functions can make it modular/easy to annotate and gather post-hoc analysis
- Hamilton pipelines specify the data movement in code -- providing a source of truth
That said, I think the ecosystem for doing this is much cleaner/easier to manage than it used to be -- the MLOps stack is far more sophisticated. Scalable/reliable compute (spark, Modin), easy storage (snowflake, new feature store technology), and more model experiment-as-a-service type systems (mlflow, model-db, etc...) have made these less of a difficult problem than it was in the past. As these permeate the industry and we develop more standards, I think that it pushes the problem up a level -- rather than figuring out exactly how to solve these, the difficult part is looping together a bunch of systems that all do it fairly well but (a) require significant expertise to manage and (b) often result in code that's super coupled to the systems themselves (making it hard to test). DAGWorks wants to decouple these from the pipeline code, enabling you to choose which systems you delegate to and not have to worry about it.
Furthermore, we think that smaller pipelines are actually super underserved in ML/data science community -- E.G. pipelines that don't have a lot of the "moving data around" problems but can be run on a single machine. I've seen these suffer from getting too complex/being difficult to manage, and we think Hamilton can solve this out of the box.
Thoughts?
Like Stefan mention in the OP, Hamilton works well with tools like Metaflow which can help with many other concerns you mentioned. How you define your data transformations for ML is an open question that Hamilton addresses neatly.
See here for an example of Metaflow+Hamilton in action: https://outerbounds.com/blog/developing-scalable-feature-eng...