We didn't do the "at least as good" thing though. In prior jobs I'd seen too many legitimate situations in which a metric declines even though there's no regression, such as bug fixes in metrics or occasional updates to the test data. Instead we committed the model evaluation to git and had to review and approve model updates.
I wish we'd done more testing for that pipeline, particularly the parts that fetched and preprocessed data. I think we had a couple bugs there over the years, or partial missing data.