Q: "are less sensors less safe/effective?"
A: "well more sensors are costly to the organization and add more tech debt so safety is orthogonal and not worth answering".
Q: "are less sensors less safe/effective?"
A: "well more sensors are costly to the organization and add more tech debt so safety is orthogonal and not worth answering".
Q: "Does [removing some sensors] make the perception problem harder, or easier?"
(note, this is literally what Lex asked, your restatement is misleading)
A: [paraphrasing] "Well more sensor diversity makes it harder to focus on the thing that I believe really moves the needle, so by narrowing the space of consideration, I think we'll get better results"
Karpathy might not be telling the truth, I don't know. But it's a much more credible pitch than you make it sound, because it's often true that you can deliver better by focusing on a smaller number of things. Engineering has always been about tradeoffs. Nobody is offering Karpathy infinite money plus infinite resources plus infinite time to do the job.
Again, I'm not saying Karpathy is honest or correct. I'm saying that the rephrasings in this comment and this thread are hilariously unfair.
Our brains do an amazing job interpreting high resolution visual data and analyzing it both spatially and temporally. Our brains then take that first analysis and apply a secondary, experiential, analysis to further interpret it into various categories relevant to the current activity.
What I’ve seen from Tesla so far indicates to me that FSD shouldn’t be enabled regardless of what sensor package they’re using, let alone based on camera data only. They need to solve their ability to accurately observes their surroundings first, especially temporally. Things shouldn’t be flashing in and out that have been clearly visible to the human eye the entire time. Additionally, this all ignores the experiential portion of driving. When most people approach something like a blind driveway or crosswalk obscured by a vehicle (a dynamic, unmapped, situation), they pay special attention and sometimes change their driving behavior.
This is true / dogma in linear / non-linear regression world, but of no real import in deep learning or Bayesian methods.
My understanding is that by focusing on fewer things (vision only), they bet to make progress faster because of the simplified organisational aspect.
In general the ‘harm to consumers’ is really just making it more likely they damage the car in a parking lot or their garage, which tells you where their priorities are (sales, Automotive gross profit). Assuming occupancy network works, the only real blind spot left is if something in front of the car changes in between it turning off and on (assuming occupancy will 'remember' the map around it when it goes to sleep).
Also, Tesla’s strategy for safety is seemingly “excel in industry standard tests, ie. IIHS and EuroNCAP”, so this might be a case of the measure becoming a target.
Meanwhile, radar is the principal sensor used in systems like automatic emergency braking across the industry. It has no intersection with any of the parking stuff because it generally has to ignore stationary objects to be useful (hence the whole "Teslas crashing full speed into stopped vehicles" thing).
That's literally trivial for a car with radar to detect.
Amazing how people talk about stuff they have no idea about when it comes to Tesla.
Now, where we disagree is you implying that cars with AEB-level radar (literally $10 off-the-shelf parts with whatever sensor fusion some MobilEye intern dreams ups) are somehow the same as self-driving cars (the goal of Tesla Autopilot).
Every serious self-driving car/tractor-trailer out there uses radar as a component of its sensor stack because Lidar and simple imaging is not sufficient.
And that's the point I was trying to make - we agree it's trivial for radar to find things they just need sensor fusion to confirm the finding and begin motion planning. This is why a real driverless car is hard despite what Elon would like you to believe. There is no one sensor that will do it. Full stop.
And this cuts to the core of why Tesla is so dangerous. They are making a car with AEB and lane-keeping and moving the goal posts to make people (you included) think that's somehow a sane approach to driverless cars.
Yet somehow, humans can drive cars with just a pair of optical sensors (mounted on a swivelling gimbal, of sorts).
In theory, a sufficiently capable AI should be able to drive a car at least as well as a human can using the same input: vision.
And Tesla lacks that, so therefore they ought not simply rely on cameras and ought use extra auxiliary systems to avoid danger to their consumers, they are not doing this because it reduces their profit margins, alas, this hn thread
Needs emphasis on can
A pair of optical sensors and a compute engine vastly superior to anything that we will have in the near future for self-driving cars.
Humans can do fine with driving on just a couple of cameras because we have an excellent mental model (at least when not distracted, tired, drunk, etc.). Cars won't have that solid of a mental model for a long, long time, so sensor superiority is a way to compensate for that.
When we look at the road, we recognize stuff in the images we get as objects, and then most of the work is done by us applying basic logic in terms of those objects - that car is off the side of the road so it's stationary; that color change is due to a police light, not a change in the composition of objects; that small blob is a normal-size far-away car, not a small and near car; that thing on the road is a shadow, not a car, since I can tell that the overpass is casting it and it aligns with other shadows.
All of these things are not relying on optics for interpreting the received image (though effects such as parallax do play a role as well, it is actually quite minimal), they are interpreting the image at a slightly higher level of abstraction by applying some assumptions and heuristics that evolution has "found".
Without these assumptions, there simply isn't enough information in an image, even with the best possible camera, to interpret the needed details.
Of course, and all this is exactly what self-driving AIs are attempting to implement. Things like object recognition and understanding basic physics are already well-solved problems. Higher-level problem-solving and reasoning about / predicting behaviour of the objects you can see is harder, but (presumably) AI will get there some day.
Basically my contention is that vision-only is being touted as the more focused path to self-driving, when in fact vision-only clearly requires a big portion at least of an AGI. I think it's pretty clear this currently means this is not a realistic path to self-driving, while other paths to self-driving using more specialized sensors seem more likely to bear fruit in the near term.
In fairness, humans have a lot more than just optical sensors at their disposal, and are pretty terrible drivers. We've added all kinds of safety features to cars and roads to try to compensate for their weaknesses, and it certainly helps, but they still make mistakes with alarming regularity, and they crash all the time.
When you have a human driver, conversations about safety and sensor information seem so straightforward. The idea of a car maker saving a buck by foregoing some tool or technology at the expense of safety is largely a non-starter.
What's weird is, with a computer driver, (which has unique advantages and disadvantages as compared to a human driver) the conversation is somehow entirely different.
This is a super important point. Whenever self-driving cars comes up in conversation it's like, "we're spending billions of dollars on self-driving cars tech, but what if we just, idk, had rails instead of roads". We're putting all the complexity on the self-driving tech, but it seems pretty clear that if we helped a little on the other end (made driving easier for computers), everything would get better a lot faster.
This is wrong and I was surprised to hear them say it was enough in the video.
We don't have car horns and sirens for your eyes. You will often hear something long before you see it. This is important for emergency vehicles. Once you hear it, a good driver will immediately slow down and pull to the side, or delay movement to give space for the vehicle.
Does this mean self driving vehicles can't detect emergency vehicles until they appear on camera? That's not encouraging.
Robotically performing an action in response to single/few stimuli with little consideration for the rest of the setting and whether other responses could yield more optimal results precludes one from ever being a "good" driver IMO.
"See lights, pull over" is not going to cut it. See any low effort "idiot drivers and emergency vehicles" type youtube compilation for examples of why these sorts of approaches fall short.
In theory, cars should be use mechanical legs instead of wheels for transportation, that's how animals do it. In theory, plane wings should flap around, that's the way birds do it. My point being: the way biology solved something may not always be the best way to do it with technology.
Wheels and legs solve different problems. Wheels aren’t very useful without perfectly smooth surfaces to run them on. If roads were a natural phenomenon that had existed millions of years ago, then isn’t it plausible that some animals might have evolved wheels to move around faster and more efficiently?
That number is probably too high for robots to do though.
Humans are weird like that.
Our eyes provide distance sensing through focusing, the difference in angle of your two eyes looking at a distant object, and other inputs, as well as having incredible range of sensitivity, including a special high contrast mode just for night driving. This incredibly, literally unmatched camera subsystem is then fed into the single best future prediction machine that has ever existed. This machine has a powerful understanding of what things are (classification) and how the world works (simulation) and even physics. This system works to predict and respond to future, currently unseen dangers, and also pick out fast moving objects.
Two off the shelf digital image sensors WILL NEVER REPLACE ALL OF THAT. There's literally not enough input. Binocular "vision" with shitty digital image sensors is not enough.
Humans are stupidly good at driving. Pretty much the only serious accidents nowadays are ones where people turn off some of their sensors (look away from the road at something else, or drugs and alcohol) or turn off their brain (distractions, drugs and alcohol, and sleeping at the wheel).
That's literally trivial for a car with radar to detect.”
That crash occurred on a car which was using radar. Automotive radar generally doesn’t help to detect stationary objects.
Further, that crash occurred on a vehicle with the original autopilot version (AP1), which was based on Mobileye technology with Tesla’s autopilot software layered on top. Detection capabilities would have been similar to any vehicle using Mobileye for AEB at the time.
Maybe it's difficult for reasons of false alarm detection (too many stationary objects that are not of interest) but you can get very good results with tracking (curious about these radars' refresh rate), STAP, and classification/identification algorithms, especially if you have a somewhat modern beamformed signal (so, some kind instant spatial information). Active-tracking can also be of help here if you can beamsteer (put more energy, more waveform diversity on the target, increase the refresh rate). Can't these radars do any of those 'state of the art 20 years ago' stuff?
There's something I don't get here and I feel I need some education...
I’m in the same boat as to not understanding why, but from what I have read the problem indeed isn’t that it doesn’t detect them, it’s that there are too many of them, and nobody has figured out how to filter out the 99+% of signals you have to ignore from the ones that may pose a risk, if it’s doable at all.
I think that at last part of the reason is that spatial resolution of radar isn’t great, making it hard to discriminate between stationary objects in your path and those close to it (parked cars, traffic signs, etc). Also, some small objects in your path that should be ignored such as soda cans with just the ‘right’ orientation can have large radar reflections.
Personally I've never seen these claims come from the mouth of an automotive radar expert, and many cars do use radar in their adaptive cruise control, so I present it as a rumour, not a fact :)
Some of the newest car radars can do some beam formimg, but not all.
Most models have multiple radars pointing in multiple directions as that's cheaper than AESA.
Only just recently have "affordable" beamformer's come to the market. And those target 5G basestations.
So the spec in most K/Ka-band models starts at 24.250GHz, where the 5G band starts. While the licence free 24GHz band that the radars use is 24.000-24.250GHz.
If this was not bad enough there has been consistent push from regulators to get the car radars on the less congested 77GHz band. And there's even less afforable beamformers for that band.
The issue is the number of false positives, stationary objects need to be filtered out. Something like a drainage grill on the street generates extremely strong returns. RADAR isn't high enough resolution to differentiate the size of something, you only have ~10 degree resolution, and after that you need to go by strength of the returned signal. So there's no way to differentiate a bridge girder or a railing or a handful of loose change on the road from a stationary vehicle. On the other hand, if you have a moving object, RADAR is really good at identifying it and doing adaptive cruise control etc.
Edit: it looks like some of the latest Bosch systems have much better performance in terms of resolution and separability: https://www.bosch-mobility-solutions.com/media/global/produc...
RADAR can have high(er) angular resolution with (e.g.) phased arrays (linear or not) and digital beamforming. I guess it's the way the industry works and it wants small cheap composable parts, but using the full width of the car for a sensor array you could get amazing angular accuracy, even with cheap simple antennas. MIMO is also supposed to give somewhat better angular accuracy, since you can perform actual monopulse angular measurement (as if you had several independent antennas). There's even recent work on instant angular speed measurement through interferometry if you have the original signals from your array.
And with the wavelengths used in car RADARs you could get far down on range resolution, especially with the recent progress on ADCs and antenna tech.
I'm not saying you're wrong, you're describing what's available today (thanks for that).
Wondering when all this (not so new) tech might trickle down to the automotive industry... And whether there's interest (looking at big fancy manufacturers forgoing radar isn't encouraging there).
So the car would be very difficult to sell since few people are willing to pay much higher insurance premiums just for that.
Which car brand do you think would take up these restrictions, and which customer is then going to buy the car with the big ugly patch on the front?
Modern phased arrays can have independent transmitters (synchronized digitally or with digital signal distribution) or you can have one 'cheap and stupid' transmitter and many receivers, doing rx beamforming, and as for complexity you mostly 'just' need to synchronize them (precisely). The receivers can then be made on the very cheap and you need some signal distribution for a central signal processor.
Non-linear or sparse arrays are also now doable (if a bit tricky to calibrate) and remove the need for complete array or rigid substrate or structure.
If you imagine the car as a multistatic many-small-antennas system there's lots that could be done. Exploding the RADAR 'box' into its parts might make it all far more interesting.
I'll admit I'm way over my head on the industrial aspects, so thanks for the reality check. Just enthusiastic, the underlying radar tech has really matured but it's not easy to use if you still think of the radar as one box.
If you wanted separated components to group together many antennas I suspect the difficulty would be accurate clock synchronization what with automotive standards for wiring. I'm still not sure I understand how they can get away without having rigid structures for the antennas, but this would be a critical requirement because automotive frames flex during normal operation.
Cars are also quite noisy RF environments due to spark plugs.
I guess what you're speaking of will be the next 10-20 years of progress for RADAR systems as the engineering problems get chipped away at one at a time.
Thanks for humouring me. RADAR is a very fun and interesting topic.
That's literally trivial for a car with radar to detect.
In principle that is correct… but radars in automotive application are unable (or rather not used) to detect non-moving targets ?Asking this because I know first hand that the adaptive cruise function in my car must have a moving vehicle in front of it for the adaptive aspect to work. It will not detect a vehicle that is already stopped.
The resolution of the radar is pretty good though, even if the vehicle in the front is just merely creeping off breaks… it does get detected if it is at or more than the “cruising distance” set up initially.
The AEB function on my car depends on the camera.
Radar is quite good at finding stationary metal objects, particularly. Putting it in a car, if anything, helps, because the station objects are more likely to be moving relative to the car...
Radar does however have the advantage of measuring object speed directly via the doppler effect, so you can filter out all stationary objects reliably, then assume that all moving objects are on the road in front of you and need to be reacted/responded to.
So I think it's the case that radar can detect stationary objects easily, but cannot determine their position enough to be useful, hence in practice stationary objects are ignored.
The car that crashed had radar, vision, USS, AND it was based on another company's technology.
The sensors are unreliable and expensive in terms of R&D. Having marginal parts which takes money from a finite R&D budget can easily result in a worse product. “They contribute noise and entropy into everything.” … “you’re investing fully into that [vision] and you can make that extremely good. You only have a finite amount of spend of focus across different facets of the system.”
His standpoint can be summed up as “I think some of the other companies are going to drop it.” Which would be really interesting if true.
"Less sensors can be more safe/effective if that allows us to focus on making effective use of the sensor information we do have, which is the result we're aiming for with this descision." would be a reasonable answer (if true), but that doesn't seem like a fair interpretation of what he actually said.
“Organizationally it can be very distracting. If all you want to get to work is vision resources are on it and you’re actually making forward progress. That is the sensor with the most bandwidth the most constraints and you’re investing fully into that and you can make that extremely good. You only have a finite amount of spend of focus across different facets of the system.”
Which was from this section: Q: “Is it more bloat in the data engine?”
“100%” (Q:“is it a distraction?”) “These sensors can change over time.” “Suddenly you need to worry about it. And they will have different distributions. They contribute noise and entropy into everything and they bloat stuff”.
Even earlier he says:
“These sensors aren’t free…” list of reasons including “you have to fuse them into the system in some way. So that like bloats the organization” “The cost is high and you’re not particularly seeing it if your just a computer vision engineer and I am just trying to improve my network.”
I listened to his answer three times and I’m not able to come up with a different interpretation than that.
Andrej eventually gets to it. But his first response was to evade. Lex is a skilled interviewer. By not letting him wriggle out of a difficult question we eventually got a substantive answer. But Andrej's first instinct was to evade. That's notable.
Otherwise I agree.
Fair enough. Seemingly practiced may be a better description. (I'm not super familiar with his work.)
This is what doesn't add up to me. Either a lot of that previous wonder-talk was actually a lie, or there's something else going on here.
Tesla dropped radar and ultrasonic due to supply shortages. Nothing to do with their AI being smart.
Many first-hand reports on Tesla fanboy forums on how no-radar and no-ultrasonic autopilot is far worse than with the sensors.
Your optical system can be good as heck till a bug hits it directly on the lense coving an important frontal area and make it behave weirdly.
In your metaphor it's like asking if you should have project managers as well as engineers on your project. And Tesla has decided that having only engineers allows them to focusing on having the best engineers. And they avoid the distraction of having to manage different types of employees.
More programmers are like having more testicles, theoretically they should enable you to have more kids, but in practice bootleneck are elsewhere.
Both your and my reasoning by comparison is equally valid
I bring up the programmers working on a project example just to illustrate how more isn't always better even if it theoretically can be.
Team focus on vision which is by far the highest accuracy and bandwidth sensor allows for a faster rate of safety innovation given a constant team size.
https://www.bloomberg.com/news/articles/2022-10-27/tesla-eng...
And the corresponding HN thread: https://news.ycombinator.com/item?id=33365065
Tesla has ca. 1000 software engineers working in various capacities. The ca. 300 that work on car firmware and autonomous driving are probably not participating in the Twitter drama.
The bottom line seems to be that the part shortages would have slowed production and cost cutting. The rest of the story seems like a fable to me. It was pretty clear Tesla removed the radar because it couldn't get enough radars.
The interview didn't really impress me. I'm sure Andrej is bound by NDA and not wanting to sour his relationship with Tesla/Elon but a lot of the answers were weak. (On Tesla and some of the other topics, like AGI).
I also expect an automated system to be better than the poor human in the drivers seat.
If you're against having multiple sensors though, the rational conclusion would be to just have one sensor, but Tesla would be the first to tell you that one of the advantages their cars have over human drivers is they have multiple cameras looking at the scene already.
You already have a sensor fusion problem. Certainly more sensors add some complexity to the problem. However, if you have one sensor that is uncertain about what it is seeing, having multiple other sensors, particularly ones with different modalities that might not have problems in the same circumstance, it sure makes it a lot easier to reliably get to a good answer in real-time. Sure, in unique circumstances, you could have increased confusion, but you're far more likely to have increased clarity.
Yes, in machine learning, pruning down to higher signal data is important, but good models are absolutely amazing at extracting meaningful information from noisy and diffuse data; it's highly unusual to find that you want to dismiss a whole domain of sensor data. In the cases where one might do that, it tends to be only AFTER achieving a successful model that you can be confident that is the right choice.
Tesla's goal is self-driving that consumers can afford, and I think in that sense they may well be making the right trade-offs, because a full sensor package would substantially add to the costs of a car. Even if you get it working, most people wouldn't be able to afford it, which means they're no closer to their goal.
However, I think for the rest of the world, the priority is something that is deemed "safe enough", and in that sense, it seems very unlikely (more specifically, we're lacking the tell tale evidence you'd want) that we're at all close to the point where you wouldn't be safer if you had a better sensor package. That means, in effect, they're effective sacrificing lives (both in terms of risk and time) in order to cut costs. Generally when companies do that, it ends in law suits.
More or less. You can take that decision on other grounds - e.g. "what would be safest to do if one of them is wrong and i don't know which one?"
The system is not making a choice between two sensors, but determining a way to act given unreliable/contradictory information. If both sensors allow for going to the emergency lane and stopping, maybe that's the best thing to do.
Predictability especially around failure cases is a very important feature. Most human drivers have no idea about the failure modes of lidar/radar.
It has everything to do with cost cutting?
Okay? Tesla is a car company and they are absolutely trying to sell a cheaper car. That's obvious to anyone that's been in one.
"Both methods have trade-offs."
Right, isn't that why most other systems use both?
He's not conspiring to trick people per se but he's also not being super clear. His position obviously makes it difficult to answer this question. It's possible he really believes this is better but if he didn't he wouldn't exactly tell us something that makes him and his previous employer look bad. Also his belief here may or may not be correct.
Is it a coincidence that the technical stance changed at the same time when part shortages meant that cars could not be built and shipped because of shortages of radars?
More likely there was some brainstorming as a result of the shortages and the decision was made at that point to pursue an idea of removing the additional sensors and shipping vehicles without those. This external constraint makes believing the claims that this is actually all around better, while hearing some reports of increases in ghost braking (anecdotes) a little difficult. Not clear if there was enough data at that time to prove this and even Andrej himself sort of acknowledges that it's worse by some small delta (but has other advantages, well shipping cars comes to mind).
So yes, sensors have to be fused, it's complicated, it's not clear what the best combination of sensors is, the software might be larger with more moving parts, the ML model might not fit, a larger team is hard to manager, entropy - whatever. Still seems suspicious. Not sure what Tesla can do at this point to erase that, they can say whatever they want, we have no way of validating that.
Here is one possible perspective from an engineering standpoint:
Same amount of $$, same amount of software complexity, same size of engineering teams, same amount of engineering hours, same amount of moving parts. One company focuses on multiple different sensors and complex fusion with some reliance on AI. Another company focuses on limited sensors and more reliance on AI. Which is better? I don't think the answer is clear.
The other point is that I am arguing that many people are over-stating the importance of the sensors. They are important, but far more important is the post-processing. Any raw sensor data is a poor actual representation of the real environment. It is not about the sensors, but about everything else. The brain or the post-sensor processing is responsible for reconstructing an approximation of the environment. We have to infer from previous learned experiences of the 3D world to successfully navigate. There is no 3D information coming in from sensors, no objects, no motion, no corners, no shadows, no faces, etc. That is all constructed later. So whoever does a better job at the post-processing will probably out perform regardless of the choice of sensors.
Lights?
Ultrasonics struggle in the rain, btw.
EDIT: Also I'm not quite positive why the image is so dark when I reverse at night. But it still is. The slope and surface of the driveway might have something to do with that... Still I wouldn't trust that camera. The ultrasonic sensors otoh seem to do a pretty good job. That's just my experience.
EDIT2: I love the Tesla btw. The ultrasonic sensors seem to work pretty reliably, they're pretty much their own system, the argument about complexity doesn't really seem to hold water and on the face of it the cameras won't easily replace them...
Now, simple color matching models are used in some fancy toasters on white bread to determine brownness. That's the most I've ever seen in appliances...
By hiding the ball that you are starting from a much more unsafe position
They are literally the least accurate of all sensors.
Radar tells you distance and velocity of each object. Lidar tells you size and distance of each object. Ultrasonic tells you distance. Cameras? They tell you nothing!
Everything has to be inferred. Have you tried image recognition algorythms? I can recognise a dog from 6 pixels, the image recognition needs hundreds, and has colossal failures.
We have no grip on the results AI will produce and no grasp on it's spectacular failures.
Driving will have to be solved without AI