From my perspective in our effort to build our machine learning platform[0], I have to look at things in terms of impact on our capacity to execute projects faster in a more repeatable way, and help our data scientist colleagues.
I keep an eye on new things, but also notice and ask them about what they use and how, then work these frameworks into our platform without compromising flexibility.
This of course comes with some frustrations when you do actual projects that involve more than one person, and where documentation that starts with "first, download your dataset to disk" becomes off-putting.
For example, we look at how to use PyTorch with S3 object storage, and you stumble on threads where technical support staff tells the asker to find examples on the internet.[1]
This brings us to try and find ways to make it work on larger datasets in the context of our platform. This is not particular to PyTorch, though the support thread is hilarious. Looking for ways to use object storage with Tensorflow wasn't obvious. The docs show how you could use files giving a path string, but you have to dig a bit deeper in the source code to become aware you could give it a file-like object, and then you have to get the bytes from somewhere, wrap it, and give it to the function that consumes data. Sure, there's the `file_io`, but again, it is not super obvious and you have to inherit that to simplify usage and reduce the "activation energy" for data scientists.
This is generally true for other parts of the pipeline, and is a reason why we don't buy into the hype of "end-to-end" machine learning or "complete lifecycle management" announcement at conferences.
For example, we do automatic model detection from code and log the models and parameters with MLflow for now so that data scientists don't have to remember or know how to. It's all done for them. MLflow has documentation on the ability to "deploy" these models. However, it breaks when these models expect higher-dimensional input (tensors) and expects a DataFrame, so we're looking into pandas' MultiIndex and things like that[2]. But this shows how something that is obvious and common in the real world lacks support, or worse, the issue is closed by an intern who doesn't see how it is a problem, which has happened, or a bot automatically closing the issue.
We reach out to people to see how they are doing things, and they reply that they write custom code to handle these cases that are not edge cases. And for us, who are working precisely to reduce "custom code" so people can train, track, deploy, monitor, and manage models consistently, reliably, and systematically, this is not good enough and drives us to solve these problems without relying on our proposed changes to be merged into the main tree or forking the repo and having to maintain that fork and conflicts.
- [0]: https://iko.ai
- [1]: https://discuss.pytorch.org/t/will-pytorch-support-cloud-sto...
- [2]: https://github.com/mlflow/mlflow/issues/3570