- simple logging of (simple) metrics during and after training
- simple logging of all arguments the model was created with
- simple logging of a textual representation of the model
- simple logging of general architecture details (number of parameters, regularisation hyperparameters, learning rate, number of epochs etc.)
- and of course checkpoints
- simple archiving of the model (and relevant data)
and all that without much (coding) overhead and only using a shared filesystem (!) And with an easy notebook integration. MLflow just has way to many unnecessary features and is unreliable and complicated. When it doesn't work it's so frustrating, it's also quite often super slow. But I always end up creating something like MLflow when working on an architecture for a long time.
EDIT: having written this...I fell like trying to write my own simple library after finishing the paper. A few ideas have already accumulated in my notes that would make my life easier.
EDIT2: I actually remember trying to use SQLite to manage my models! But the server I worked on was locked down and going through the process to get somebody to install me SQLite was just not worth it. It's also was not available on the cluster for big experiments, where it would be even more work to get it, so I gave up on the idea of trying SQLite.