“The latter being when the training or test data follows a different distribution to the in-operation data”
This form of ML bug is the most challenging to catch. The true in-operation distribution is often unknown which makes testing for such bugs a very challenging problem. Any thoughts on this?