Obviously they tried it multiple ways have brought receipts, but nonetheless it seems surprising that it wouldn't be of benefit to bring as much of that kind of metadata to the model as possible. You'd think depth and segmentation in particular would basically just be a straight shortcut without which the model spends its own time and effort re-deriving that stuff.
I'd also be interested in how post-processing fits in with this. Like if you've got weather effects, film grain, tone mapping, etc, I would have thought the model would do better working on the image before those processes.