I've been working in AI/ML and some CV projects over the last decade.
I've seen far too many cases where the "algorithms" and modeling teams had no concern or even a concept for how the input systems for data that would be used to train models, and later inference, mattered to the quality of outcomes.
In CV and computational photography cases, there was little concern or understanding for how photography and imaging actually works, nor consideration for how variances in hardware, configuration, or the capture pipeline in general might affect those models when doing training data capture in parallel. Adding another layer, consideration for how variances between the capture pipeline and actual inference pipelines might need to be accounted for when designing the overall system and how to approach data collection and curation for training + evaluation, as well as to build-in robustness. (similar thoughts apply to concepts like bias and fairness in models)
Example: training vision models using one set of imaging hardware and configurations while applying those models on very different imaging hardware with different characteristics.
To summarize the above, not caring about calibration in CV is like not caring about how variances in feature extraction/embedding generation will affect the overall quality of your results.