> Data scientist here, stuck in the Dark Ages of "deploying" my models by writing bespoke Python apps that run on some kind of cloud container host like ECS. Dump the outputs to blob storage and slurp them back into the data warehouse nightly using Airflow. Lots of manual fussing around.
Oh, man, been there! So, first, I want to say there are a lot of ML/data platforms -- largely because there are so many problems to solve, and they're not one-size-fits-all solutions.
> What the heck are all these ML and data platforms, how do they benefit me, and how do I evaluate the gazillion options that seem to be out there?
You probably don't need all of them, or all that many. As in everything, it depends on your pain-points. Given that you have a lot of manual fussing around, you probably want something to reduce it. We've found airflow to be painful, but a lot of dev teams/DS have airflow already integrated into their platform, so we wanted to build something that allows data scientists to plug into it. So for people who don't like airflow, the idea is that you could express your dataflow in Hamilton and DAGWorks can ship it to airflow (or any other orchestration system).
> For example, I recently came across DStack (https://dstack.ai/) and have had an open browser tab sitting around waiting for me to figure out WTF it even does. DAGWorks seems like it does something similar. Is that true? Are these tools even comparable? How would I choose one or the other? Is there overlap with MLFlow?
DStack is definitely a different approach, similar space. Hamilton is organized around python functions and dstack is more of a high-level workflow spec (reminds me of something we had at my old company). So Hamilton can model the "micro" of your workflow, whereas DStack models the "macro" -- managing artifacts. DStack could easily run Hamilton functions. What nothing we've found out there do (except perhaps kedro) is model the "micro" -- E.G. the specific fine-grained dependencies so you can take a look at your code and figure out how exactly it works.
Re: MLFlow -- DAGWorks + MLFlow are pretty natural connectors. Hamilton functions can produce a model that DAGWorks would be able to save to mlflow. DAGWorks is more on the data transform side, and doesn't explicitly say how to represent a model.