This never made sense to me. You certainly need enough data, but how you interpret and process that data is far more important.
This never made sense to me. You certainly need enough data, but how you interpret and process that data is far more important.
Waymo's autonomous platform is a frankenstein of various machine learning techniques, and much of it isn't glamorous, it's less contigent on big breakthroughs than it is on elbow grease. Google demoed as proof-of concept full autonomy in 2012, and much of what they've been doing in the 4 years between then and now is the tedious job of addressing and validating their system across the full spectrum of edge cases that must be dealt with if they ever hope to foist their safety critical software upon the public.
It's not clear to me that Tesla's current development paradigm will ever be sufficient to completely take the human out of the loop. Tesla's approach is incremental, and I suspect they'll have to make some big changes if they wish to fully close the gap. Waymo has kept their eye on the prize from day 1.
Think of it this way: if Tesla wants to test a particular algorithm for a particular driving situation, they can "play it back" over an enormous amount of real-world situations. They will have tons more potential edge cases with which they can validate their algorithms.
Big data is no where near as much a competitive advantage as it was three years ago. It seems not everyone outside the field has noticed that though.
Image classification seems like it would be very different, most importantly that 99.9% "correct" would be a great achievement, but for self-driving cars a .1% failure rate would be completely unacceptable.
Please oh wise ones how do we simulate nlp data, numeric data, finance data, biological data and anything else machine learning is used for.
Oh you are able to classify dogs and cats in images after a 2 hour youtube. How nice.
Renesd is correct that "big data" is overblown. There are diminishing marginal returns - you need orders of magnitude more data for the same incremental gain (and this blows up well beyond however millions of cars Tesla can hope to run).
You're correct that data augmentation is only a marginal technique to squeeze out more performance, and not generally possible in many domains.
Sure, they can push a beta algorithm to cars and record high-level decision making between human & algo, verifying it's not totally out of whack. But that's hardly something that is going as training data into the models.
Big public opinion perspective here too.