As I recall, one group had fairly good success, but eventually someone figured out that their data set had images from a low-COVID hospital and a high-COVID hospital, and the lettering on the images used different fonts. The ML model was detecting the font, not the COVID.
[a bit of googling later...]
Here's a link to what I think was the debunking study: https://www.nature.com/articles/s42256-021-00338-7
If you're not at a university, try searching for "AI for radiographic COVID-19 detection selects shortcuts over signal" and you'll probably be able to find an open-access copy.