During my PhD this issue came up amongst those in the group looking into compressed sensing in MRI. Many reconstruction methods (AI being a modern variant) work well because a best guess is visually plausible. These kinds of methods fall apart when visually plausible and "true" are different in a meaningful way. The simplest examples here being the numbers in scanned documents, or in the MRI case, areas of the brain where "normal brain tissue" was on average more plausible than "tumor".
[1]: http://www.dkriesel.com/en/blog/2013/0802_xerox-workcentres_...