Hi, I am one of the developers. This makes total sense, we were thinking of doing alerts. Do you have any specific alerts in mind? Do you think they should be customizable or general anomalies be enough?
I don't think you are missing a better workflow. Our current use case is monitoring the hardware while monitoring model training. In this case researchers tend to take a look at how their models are doing from time to time and they could check on hardware at the same time. This wouldn't be the case for general servers.