Can you elaborate on what you mean here? Do you mean that Tesla will only be using Computer Vision for autonomy? Doesn't or wouldn't Lidar compliment CV? Or is there a practical design consideration that makes them mutual exclusive?
Can you elaborate on what you mean here? Do you mean that Tesla will only be using Computer Vision for autonomy? Doesn't or wouldn't Lidar compliment CV? Or is there a practical design consideration that makes them mutual exclusive?
Humans have far higher resolution sensors and the most advanced computer ever seen behind them. And even that fails at a far higher rate than we want self driving cars to be.
The thing we most bring to driving is navigating complex, low speed environment changes - but we're not great at that either (see the number of toddlers run over in their drive ways, for example).
https://crashstats.nhtsa.dot.gov/Api/Public/ViewPublication/...
Thanks, but did you mean to say that people fail at a far lower rate that we want self driving cars to have? Maybe I misunderstood your point?
Expanded version of the argument: "Karpathy says 'people drive vision-only' and apparently intends that to convince us that vision-only is good enough. But (1) those people driving vision-only are using human brains, whose abilities we have not yet come close to duplicating, and (2) even with the astonishing abilities of the human brain, those people driving vision-only make a lot of mistakes and have a lot of accidents, and we want self-driving cars to be much safer than that. So the fact that people drive vision-only is no reason to think that self-driving cars should do likewise. They're trying to be safer than people, with less computing power; why shouldn't we make up for that by giving them extra sensors?"
This is really not true. The phone in your pocket is just plain better. More total pixels per unit area, much greater dynamic range, able to sample over time periods an order of magnitude faster.
The reason your vision seems better is because your brain is amazingly good at synthesizing a picture of the world around you. But all that data (the sphere around you is something like 150 MP at eye resolution) is an illusion. You're only looking at a million pixel(-equivalents) at once, thereabouts.
[1] Your eyes come close only in the middle half a degree of your fovea. Everywhere else your brain gets a blurry mess and has to extrapolate.
Your statement on phone cameras being better than the human eye isn't true today, though.
The human eye has FAR better dynamic range than any camera built to date, at any price point. This matters a great deal on sunny, cloudless days in a city with tall buildings casting dark shadows, for example.
Just like how the biological system of "stupid sensor, smart thinking" helps biological organisms, that's what is going to need to be done for computer vision as well.
About 1/3 of the human brain is dedicated to vision processing in some way. Think about that a second. One THIRD of the best, most powerful organic computer known is required for us to see what we see and we are still fooled by the things human eyes can be tricked by. It's going to take a lot of neural network training to duplicate that. Fortunately, the vision skills required to drive are a subset of overall vision capabilities.
Not at all true. All you need to do is take your camera into a dark room for proof. It can take useful pictures in environments you can't see. And with some manual control and safety precautions (seriously, don't actually do this) you'll note you can shoot useful photos of things that are very near bright sources like the sun where your eyes would be completely useless (and irrevocably damaged, again do not try this).
What you're complaining about isn't dynamic range, it's exposure control. A human brain, again, makes much better decisions about what parts of the environment are "important" when setting the aperture (iris). So the stuff you want to see is visible, where the camera is mostly just going to guess that the center of the frame is what you want and will routinely leave stuff over or underexposed.
Your last paragraph has the correct analysis but the wrong conclusion. It's the vision processing that makes the difference. The optical systems of a semiconductor camera absolutely are better, so a control system based on optics can absolutely be better.
No long exposures, no multiple exposures with varying exposure -- a single frame. I don't know of any 24-32-bit per channel sensors, and that's what you'd need.
We agree on vision processing. I thought I made it clear that it's a strong "backend" to vision (the brain or compute behind it) that makes it work so well, and that will be the case with good self-driving cars as well. I must have misstated something.
Vertebrate eyes don't have 24 bits of sensitivity, that's just insane. Rod cells are neurons, they either fire or they don't, can fire at most at about 10 Hz, and you have about a million of them at most across your whole field of vision. Do the math. What you say isn't even physically possible.
10Hz? Simply not true. Fighter pilots can identify aircraft shown on a screen for 1/250th of one second. Regular schmoes can see the difference between 30Hz and 60Hz video easily.
I am not arguing about pixel count of a camera vs the center of vision of human eyes.
I'm saying that... Nevermind. I've already said it several times and apparently you have a PhD in all things vision.
Dynamic range doesn't require bit depth. You can have an 8 bit sensor with 20 stops of dynamic range. You'll lose color/luminance resolution of course but as long as 255 is capturing a light that's 2^20 times stronger than 0 that's 20 stops of dynamic range.
But more than that theoretical point there are already sensors pushing well beyond what humans can do in terms of dynamic range. Apparently humans have a respectable 10 stops and sensors are already in the 15 to 20 stop range and pushing beyond it:
https://www.eoshd.com/2014/11/new-sony-sensor-21-stops-dynam...
To map it into 16bit values you can just use a curve to distribute the bit depth unevenly across the dynamic range. Older DSLRs did that to get perfectly usable images with just 10 bits per channel.
It's not even a secret. Light capture in a phototransistor[1] is a basically direct process without a whole lot of loss or wasted surface area. Those cone cells are living things and only have a little volume to dedicate to pigment chemistry.
[1] CCD's of course can do an order of magnitude better still, and those are easily cheap enough to put in a car.
Modern consumer cameras have far surpassed what our eyes can do in all the metrics you describe (dynamic range, resolution, response rate) and many others. It's not even close.