As someone who wrote a Deep Learning-based anomaly detection pipeline used in real world I doubt they know precisely what the issue was. It could have been a simple flagging slightly over whatever threshold they set to their system, maybe ticking off a tiny majority voting of their "suspicion" detectors. Not having humans involved, nor having clearly human-interpretable results will be always an issue at scale.