The idea that you can throw a cheap camera in a car and think you can achieve the same always seemed strange to me.
The idea that you can throw a cheap camera in a car and think you can achieve the same always seemed strange to me.
Unfortunately resolution of your eyes drops off quickly from the center of vision. Yes you can move your eyes to focus on different things, but so can cameras.
Sure the mental image you build is high resolution, but not that directly related to reality. A good example of this is the numerous optical illusions that depend on you looking at a point of an image, building a mental representation of that image, then finding a conflict whenever you move your eye.
Link that seems to be the source of the 576MP number: https://art-sheep.com/the-resolution-of-the-human-eye-is-576...
Not what I'd call a particularly scientific explanation.
True, but we can move our eyes around very quickly. Cameras in self-driving cars are usually fixed.
Also, human eyes have better dynamic range.
Let's not forget about human experience. It gives context to any visual perception, and can't be replicated by cars. We simply have access to a larger, more complex environment, and the environment fosters our intelligence. It's easy to overemphasise the brain to the detriment of environment, but in reality the environment is the main cause for the development of intelligence in the brain. The corollary is also true - a lower complexity environment will lead to lesser intelligence, no matter how powerful the brain is. Intelligence springs where agent needs meet with external limitations, those limitations guide its development.
By analogy, the environment is like the training set and the brain is like a deep neural net. A deep neural net with a poor training set is not going to be accurate. That's why DeepMind, OpenAI and others have focused so much on artificial environments (games) to train agents - they know that environments are the key to training advanced AI. A mere dataset is like a static environment where nothing new happens. Research is transitioning from datasets to simulated environments for the next step in AI. I'd even go as far as to track the evolution of AI by the evolution of environments and environment models.
As for the point about the training environment, a lot of what AI cars do is learn from simulations.
unless they wrote it down.
You can read all the books and notes about WW1 battles, you'll never come close to feel what they felt in the trenches.
Even if that's true, I (and almost anyone) can also drive cars in computer games at 1080p or less.
This is a spectacularly misleading claim. Humans have a very small field of high-definition vision, and a wide range of low definition vision, easily exceeded by almost any camera system. Where humans excel is in the heuristics of taking marginal information and teasing out meaning (not always the correct meaning), though neural networks approximate that system.
Tesla has something like 10 cameras running 1280x960 @ fps or similar. With so few data points (and no color) you can't identify similar vehicles. After all snow plows, school busses, ambulances, bikes, motorcycles, police, etc all have different behaviors. Some even signal cars with spot lights, police red/blue lights, brake lights, blinking high beams, etc. All would be invisible to lidar and contribute to lidar controlled cars acting less like human driven cars.
Biology got to be just good enough to make it to tomorrow. Photosynthesis is one of the oldest mechanisms around right? Quick google and at the moment that turns sunlight to energy tops out at 6%, now look at solar panels https://upload.wikimedia.org/wikipedia/commons/5/5d/PVeff%28...
Now, commenting on biology and electronics are topics vastly outside of my ability. I'm just trying to make the point that we can't be the best it can get, surely? We long left the time where industrial progress meaning "pick one things and do it well"
It's all about the data processing and the breakthrough that will enable human-ish self-driving cars will be there and not in more accurate sensors.
yep, resulting rich details simplify stereo matching a lot. Instead of complicated algorithm it allows our brain to just run simple displacement match in extremely parallel fashion natural to the brain.
>The idea that you can throw a cheap camera in a car and think you can achieve the same always seemed strange to me.
it isn't strange, it is just a bit early. Running 1MP stereo on a single core P4 in 2004 was just pitiful. These days a 20MP sensor costs under $30 and 16-32 cores is having much better time with it (using GPU is much cheaper and several times faster - it is just that i have these workstations around with good CPUs and no GPU to speak of. Tesla for example does run powerful GPU for their cameras.) So, scroll forward another 15 years, and we'll have nice stereo with something like 200MP sensors. I honestly don't see how lidar can do even just 4MP at minimally usable 30FPS in any foreseeable future (as probing at 150m is 1 microsecond per pixel).
However, the focal point with respect to the sensor is indeed miscalibrated in some units, and good units can become miscalibrated over time. There are some after-market correction devices available.
:)