Self-driving is way harder than I thought, says Elon Musk
youtube.com
youtube.com
Right now, the camera app on the Tesla will only show the side cameras and the rear camera when actively driving. I’ve wondered, if they showed the front facing camera, would I be able to drive by looking solely at the camera output? Is the resolution sufficient, and the dynamic range sufficient for me to drive safely at all times that I can with my human eyes looking out the window? I suspect not.
And yes, the dynamic range is a huge issue. I’ve had my Tesla discontinue autopilot when driving a curvy mountain road at sunrise/sunset, as the car enters and exits areas with direct sun at a low angle. I’ve wondered how Tesla ever plans to achieve their self driving goals with the currently deployed hardware if it can’t pass that simple test. It wasn’t discontinuing autopilot because the software failed, but the sensors appeared incapable. As a human, the driving situation sucked, but I, and every other driver on this stretch of road at that time, were able to continue driving, with just our eyeballs as input. (I should have forced a camera save, so I could analyze the camera feeds later)
It seems like someone could simply prove there exists these specific real world scenarios, which Tesla cannot and will not be able to solve, in order to launch a class action lawsuit. In theory, Tesla could counter that they’ll provide a camera upgrade to resolve the limitations… and the charade will continue.
First, you have to manufacture to automotive grade standards. So that right there is going to set you back a generation or two. You can't just pick some consumer grade sensor off the shelf.
Second, you now have to confirm to ASIL safety standards. A lot of the work we did on these sensors was to build in continuous self testing circuitry. For every frame the sensor would actually inject known values into the physical pixels, read them out and process them like a real picture, and compare with the reference before sending data out to the ISP.
I'd also like to know what webcams that aren't insanely expensive have better dynamic range. Is that really true? You list this as the third spec, but this was the most important spec to design for after safety. In part because dynamic range is also important for safety. And it has to be high dynamic range without too many artifacts. The sensor I worked on could do three exposures simultaneously on the same sensor and could either stitch them together on-die or stream out all three exposures to an external ISP. I've never heard of a consumer camera sensor with that kind of capability, but then I don't follow the market much anymore.
Sensitivity is obviously also important. I think the bayer pattern was RGBW and sacrificed color accuracy for sensitivity. Higher resolution would compromise sensitivity for a given die size.
Resolution and frame rate could be better, but frankly it's about good enough for what they're trying to do.
https://en.wikipedia.org/wiki/Image_processor
Much of the advancement in modern digital cameras is at the ISP layer, for example Apple invest massively here each year to improve the iPhone’s cameras.
At the very least, the sensors on the Tesla appear to have no IR filter on key cameras (orange cast on side cameras etc) to improve low light performance. Virtually al consumer grade webcams will have an IR filter, so it’s not quite grab a crappy webcam and shove in the car as some comments suggest!
Tesla do also have a custom ISP, it’s part of the custom silicon chip in their standard fit “FSD computer” stack:
If the sensors on the Tesla were really that good, why did they blame them on the accident where it could not distinguish between a white tractor trailer and the horizon? In that case, plenty of light was available so it points to saturation or low dynamic range. If the sensor is saturated and has no way to physically limit the amount of light incident on it, there is no amount of clever software or filtering that can provide anything other than noise, but I’m sure you know this. It’s also a function of heat, all that light heats up the sensor and has to be removed and dissipated some how (and quickly) before the next reading is taken. The more heat, the more noise and depending on the noise process model, the less overall range. High sensitivity and high range are conflicting goals and subject to design trade offs that when both are necessary increase the overall system complexity.
The whole “humans can drive via vision” thing is a misnomer because human eyes are extraordinarily sensitive, ultra wide angle, stabilized (basically like having a gimbal), accurately temperature controlled via your body, and don’t have a fixed iris. It seems that in order to make that claim, you’d need to first develop a sensor that at least meets or exceeds the specs of the human eye, which no one has really done yet, at least not in an economically and reliable enough package to put on a car.
Given enough time, money and equipment it could be done but the sensor would look (and cost) more like something you’d see on an expensive astronomer’s telescope.
One difference between true intelligence and what we have in NNs today is, if you showed an image or two of an object to a person they could identify that object in other orientations and lighting conditions pretty well. To do the same with NNs requires thousands of examples and counter examples to train. This is why with driving how could you have enough examples of strange issues to ensure the car responds safely?
The only advantage I see today for cars driving is they don't get bored or distracted or drunk, but they certainly are not intelligent.
There’s work from 2019 on gauge equivariant NNs that looks to solve this, there’s a paper and explainer video from Qualcomm Research about it. Also, Tesla presumably does have thousands of examples for most driving scenarios. There’s also lots of work in the ML community on generalization—that’s arguably what all the research is about in some way or another. I don’t intend to be flippant, but I would bet that these are issues that the researchers at Tesla are well aware of.
And this man according to his fans is going to save the planet? What a joke.
Maybe that's just my ideologic biases, but I'm quite disappointed in Elon's ventures
All updated in real time, as a background process once you get experienced enough... sometimes you forget where you're going, and end up at Work, or some other place. ;-)
---
Elon doesn't seem to have an adversarial team working against the AI. He should. This would help make up for the fact that situations requiring evasive action are incredibly infrequent, and thus underweighted in the training data by orders of magnitude.
The amount that the Tesla Autopilot engineering team has accomplished is really impressive, and I don't think it's worth giving up hope just yet.
I wouldn't anticipate Tesla shipping with worldwide FSD capability in the next 10 years, but I think within 20 to 40 years, this approach to self driving cars will be the most reliable and flexible, assuming development continues.
In terms of delivering a product, many LIDAR based self driving solutions will absolutely beat a fully AI driven design to market. However, the potential for adaptability in the Computer Vision approach is much greater than current alternatives.
Overall, I don't think it's worth dismissing Tesla entirely, but it's also important to keep in mind the long timeline for the development of adaquate computer vision technologies.