A lot probably comes down to the size of the moat. If some fairly modest dataset is "good enough" that implies needed data will be widely available can be used by anyone for their automation algorithms. (Regulation could also force some level of standardization at this layer.)
Alternatively, truly vast data sets could turn out to be the difference between decent assistive driving systems, autonomy on highways, and more broadly useful and enabling self-driving technology. Such data sets could end up being outside the capabilities of a few companies who got there first.