That was Adam years ago in the Self-driving world. Now, it costs only few hundred dollars only, starting with $99.
Resolving that correctly takes time (in ms), adds complexity and will sometimes be incorrectly judged.
Since the visual data is the more accurate the vast majority of the time, it will anyways take precedence over the other input. As humans have proven that visual is technically enough, they decided it makes more sense to squeeze the most out of the visual, rather than collecting other data, crunching it, then (in most cases) discarding it.
I am not sure they are right, and am pretty sure that even if so - they need better cameras.
But misquoting them doesn't really help your argument.