Does anyone have insights on what QA looks like at Tesla for FSD work? Because all of these seem table-stakes before even thinking about releasing the BETA FSD.
Does anyone have insights on what QA looks like at Tesla for FSD work? Because all of these seem table-stakes before even thinking about releasing the BETA FSD.
Tesla is not exactly in love with QA. Especially for FSD.
FSD is mainly 2 things: 1. (By far most important) shareholder value creating promise, that's been solved for 6 years according to their CEO. 2. Software engineering research project
What FSD is not is a safety critical systems (which it should be). They focus on cool ML stuff and getting features, with any disregard for how to design, build and test safety critical systems. Validation and QA is basically non-existent.
Based on there presentation, they for sure have a whole load of tests, many built directly from real world situation that the car has to handle. They simulate sensor input based on the simulation and check the car does the right thing.
They very likely have some internal test drivers and before the software goes public it goes to the cars of the engineers.
Those are just some of things we know about.
I have no source on their approach to testing safety critical systems, but we do know that they have a lot of software that has based all test by all the major governments. They are one of the few (or only) car maker fully compliant to a number of standards on automated breaking in the US. We have many real world example of videos where other cars would have killed somebody and the Tesla stopped based on image recognition.
So they do clearly have some idea of how to do this stuff.
So when making these claims I would like to know what they are based on. It might very well be true that their processes are insufficient but I would actual know some real data. Part of what a government could do, is forcing car maker to open their QA processes.
Or the government could (should) have its own open test suit that a car needs to be able to handle, but clearly we are not there yet.
I think you and I must've watched a different video.
1. I know people working at Tesla.
2. Much more important one - Elon's Twitter feed. They're doing last minute changes, and once it compiles and passes some automated tests, it's tested internally only over few days before it's released to the customers. Even if they had world class internal testing (they don't), for something having to work in such a diverse environment like self driving system without any geo-fencing, those timelines are all you need to know.
That's why I bought/will keep buying Toyota/Lexus.
https://www.euroncap.com/en/results/tesla/model+y/46618
Same for NHTSA:
https://www.nhtsa.gov/vehicle/2022/TESLA/MODEL%2525203/4%252...
Because of the FSD false promises, Tesla encourages dangerous behavior from drivers.
I don't want to be next to a Tesla driving in autonomous mode while it's driver at the wheel is not paying attention to me.
Yeah, they have QA. But for the problem they claim they’re solving (robotaxis) and speed of pushing stuff to customers (on the order of days) it vastly, vastly insufficient. And it lacks any safety lifecycle process regards - again, just look at the timelines. Even if you’re super efficient, you cannot possibly claim you can even such a basic things like proper change management (no, commit message isn’t that) or validation.
well, if you don't get the software pushed to the QA team (the customers), how else are they going to get it tested?
tesla does extensive, meaningful vetting of these updates. i'll let you do the research yourself so that maybe you can quit spreading misinformation
It is well known that issues happen despite extensive vetting. The presence of vetting does not exclude the absence of issues. For example, lets say there are 1000 defects in existence and your diagnostic finds 99% of them. So it finds 990 defects. So now there are 10 defects remaining that are not found. Next you have another detector that finds all defects. How many defect does it see? 10. It is usually going to see about ten defects. Your expectation given vetting is taking place should that be you will tend to observe about ten defects in this particular situation.
So lets say you are someone who is watching that second detector. You observe that you keep seeing defects. Probably, because of selection bias on results, you observe this with an alarming 100% rate of finding defects - the few times no defects happen it doesn't get shared for much the same reason we don't pay attention to drying paint. Can you therefore claim there is no diagnostic that is detecting and removing defects?
Well, not really, because you aren't observing the defects unconditional on the process that removed them. Basically, when you observe 10 defects, that doesn't mean there weren't 990 other defects that were removed. It just means you observed 10 defects.
So the actual evidence you need to use to decide whether there is a diagnostic taking place is evidence that is more strongly conditional on whether it is taking place. You need something that can observe whether the 990 are being caught.
In this case, we have video evidence of extensive QA processes. So that is much stronger than the evidence that defects show up, because defects show up both when we do have a vetting and we don't have a vetting. And each reasonable test case ought to be decreasing the chance of a particular defect that the test case is testing for.
For a much more thorough treatment on this subject check out:
https://www.lesswrong.com/tag/bayesian-probability#:~:text=B....
So that is basically why he disagrees with you.
Tesla definitely needs to root cause the defects that were found and make improvements on the existing vetting process but it is very obvious they do have these processes.
I do have high hopes that the work they've done on the FSD stack will make for a significant improvement to basic AP whenever they get merged (assuming it ever happens; it has been talked about for years). That'd be nice.
completely demonstrably false
> speed of pushing stuff to customers (on the order of days)
this is also false and doesn't happen
> you cannot possibly claim you can even such a basic things like proper change management (no, commit message isn’t that) or validation.
you know absolutely nothing about the internal timelines of developments and deployments at tesla and to suggest it's impossible without that knowledge is just dishonest
Head of AP, testified under oath, that they don't know what's Operational Design Domain. I'll just leave it at that.
> > speed of pushing stuff to customers (on the order of days) > this is also false and doesn't happen
Never ever Musk tweeted about .1 fixing some critical issues coming in next few days? I must live in a different timeline.
> > you cannot possibly claim you can even such a basic things like proper change management (no, commit message isn’t that) or validation. > you know absolutely nothing about the internal timelines of developments and deployments at tesla and to suggest it's impossible without that knowledge is just dishonest
Let's assume I have no internal information. If it looks like a duck, swims like a duck, and quacks like a duck, then it probably is a duck.
Lets say you have a baby which is being born. You tweet, birth in ten days! You can't then say, look: here is a tweet, it proves that baby development in tweeters actually has a several day lifecycle and moreover it proves that baby having mother's don't do proper pre-birth routines, because the tweet isn't the process that created the baby.
It is separate.
> If it looks like a duck, swims like a duck, and quacks like a duck, then it probably is a duck.
Right. So the fact that we have video evidence of internal processes, including QA processes, is much more like looking, swimming, and quacking like a duck much like having video evidence of a mother with a round belly for months before the tweet would be evidence that the babies don't take weeks to develop.
So when Elon has also tweeted that a launch was delayed, because of issues that were discovered - which does happen, as I'm sure you know if you follow his tweets as you imply - then that would be evidence congruent with the video evidence we have of QA processes existing within the company.
For example, did you predict, based on the speculation of Tesla being incompetent with regard to safety, that they have the lowest probability of injury scores of any car manufacturer? Because they do.
Did you predict, based on speculation about Elon Musk's incompetence in predicting that self-driving would happen, that there are millions of self-driving miles each quarter? Because there are.
Did you predict, based on speculation about Tesla incompetence in full self-driving, that the probability of accident per mile is lower rather than higher in cars that have self-driving capabilities? Because they do.
I know this sort of view is very controversial on Hacker News, but I still think it is worth stating, because I think people are actually advocating for policies which kill people because they don't actually know the data disagrees with their assumptions.
I'm not sorry that annoys you, because it shouldn't.
Oh the data is great. I like the data. I'd take the data out to dinner. It's completely besides my point, and you continuing to be obtuse and rephrasing things this way, is not only a strawman, but it's rude.
> Your views probably do have evidence that supports them. I want to see the evidence you are using, because I think that is important.
Not every policy decision is driven by data. Some are driven by reasoning and sensibility, as well as deference to previous practices. So your whole data-driven shtick is just that... a shtick.
You claim that I said that we should focus on the good, but I didn't. I claimed that we should look at the data.
Now I feel as if you are trying to argue that looking at data is wrong because not all decisions should be made on the basis of the data. This seems inconsistent to me with your previous assertion that my ignoring data was bad, because now you argue against your own previous position.
That said, uh, datasets related to bayesian priors support your assertions about deference in decision making. So you could, if you cared to, support that claim with data. It would contradict your broader point that I should not want to have data, but you could support it with data and I would agree with you, because contrary to your assertion I was making an argument for evidence informed statements. Your inference about whether I think the evidence leans should not be taken as an argument that I believe my positions would always be what was reached by looking at the evidence, because I don't think that is true. I'm obviously going to be wrong often. Everyone is.
Unfortunately, I think you lie too much and shift your goalposts too much. So I'm not going to talk to you anymore.
I also caught you editing out what was an excessively rude comment. I'm gonna pass on further conversation, thanks.
No, you didn't. You said please ignore all the times when I'm wrong in favor of when I'm right. This is what you actually said.
You seem to be using language incorrectly. You seem to me to be confusing "said" with "meant" and in this case it seems that what you "said" was very different from what you "meant" so much so that I'm strongly getting the impression that you are lying to me, but if you aren't - then it is because confusion on the difference between said and meant.
Words have meanings. They have meanings independent of your own desires, so whatever you meant to say - it doesn't matter at all - that is not what you said. Please ignore is fairly characterized as a request for ignoring things, because it maps to something like it would be pleasing if there was ignorance. Notice it is you who is claiming of me that it would be pleasing to me if there was ignorance. I'm not making such a claim - you said this - maybe you didn't mean this, but you absolutely said it.
I'm not unfairly characterizing your words: this is the actual implication of the words you used, because it is the implication of the meaning of the words - maybe it isn't the meaning you desired, not what you meant, but it is the meaning of the words.
To kind of highlight how extreme what you claim you said is from what you actually said is, notice that you imply belief about me when talking about me requesting ignorance. Yet now, in your claim about what you actually said, you imply a position you hold: that data isn't dispositive. You don't even have the same reference in what you said versus what you now claim to have said. You did not say what you claimed to say. If you meant that, you should have said that, but these things are very very far diverged.
If you want to state, of your own belief, that data isn't dispositive, you should state that. Instead you consistently referred to a false reference to my own beliefs. Notice, even when you tried to correct my interpretation, you did not switch the reference class to your own belief, you said things like "your argument is ridiculous" which is still talking about me - not your own position that data is not dispositive. So the reference class which is not targeting me, but the general properties you now claim to have said, is not truly there.
> I also caught you editing out what was an excessively rude comment. I'm gonna pass on further conversation, thanks.
As you can imagine, with your confusion about meant versus said, I've been finding you to be lying about both your own views and mine. So if I seem a bit rude, it is because I'm kind of assuming you are smart enough to already realize all these things. I don't mean to assume you are so hostile as to know all this, but then pretend not to, but it is just one of the valid explanations for your behavior. I tried to edit my comment to remove my frustration and I'm sorry you had to see it like that.
> incredibly obtuse manner.
Which leads to this. You stated that you think I'm being obtuse, but my first assumption was that you weren't just directly lying about what I was saying. Your claim, interpreted in the way you said it, not the way you meant it, is a lie about my belief. So I tried to make your words have a meaning that would make them true, not false. I tried to be charitable, but was confused, because it really seemed like you were lying about my beliefs given your statement. This wasn't me being obtuse. This was me trying to be charitable, but being very confused, because interpreted according to what you actually said - not what you meant - you were strictly speaking stating falsehood about my beliefs.
Also, none of that is self driving. This data talks about AP, not FSD. FSD is also not self driving by any means (it's level 2 driver assist), but that's a detail at this point.
For example, elsewhere in this comment thread, someone threw out a random statistic of 400:1 as part of their argument, but this seems to me to be something like six orders of magnitude diverged from a data informed estimate.
To try and contextualize how big an error that is - it is like thinking that a house in the Bay Area has the same cost as a soft drink.
I think if we have to cite our data we are less likely to do that sort of error and more likely to catch it when it is done.
I definitely don't think FSD is magically safe. So if you think that is what I'm trying to say, please update your beliefs according to my correction that I do not believe this. I think anyone driving in FSD should remain vigilant, because it can make worse decisions than a human would.
1) those are statistics for the old version, the new version might be completely different. I've had enough one-line fixes break entire features I was not aware of that my view is that any change invalidates all the tests. (Including the tests that Tesla should have but doesn't) Now probably a given update does not cause changes outside its local area, but I can't rely on that until it's been tested.
2) the self-driving is presumably preferentially enabled for highway driving, which I assume has fewer accidents per mile than city driving, so comparing FSD miles to all miles is probably not statistically valid.
Just for context - I've been in a self-driving vehicle. Anecdotally, someone slammed on the breaks. The car stopped for me, but I was shocked: for hours before this the traffic hadn't changed, it was a cross country trip. I think I would have probably gotten in an accident there. Also anecdotally, there are times where I felt the car was not driving properly. So I took over. I think it could have gotten into an accident. Basically, for me, the best explanation I have for the data I've seen right now is that human + self-driving is currently better than human and currently better than self-driving. The interesting thing about this explanation is how well it tracks with other times where we've technology like this before. In chess playing for example, there was a period before complete AI supremacy (which is what we have now) where human + AI was better than AI.
I like the idea of being safe, so if the evidence goes the other way, advocating for only humans or only AI doing the driving, I want to follow that evidence. Right now I think it shows the mixed strategy is best and that is kind of nice to me because it implies that the policy that best collects data to reduce future accidents through learning happens to be the policy that is currently being used. I like that.
The probability of an accident for any driver assistance system will ALWAYS be lower than a human driver - but that doesn't mean the system is safe for use with the general public!
People like me are not advocating for "killing people" because we aren't looking at data - it's that no company has the right to make these tradeoffs without the permission and consent of the public.
Also if this was about safety and not just a bunch of dudes who think they are cool because their Tesla can kinda drive itself, why does "FSD" cost $16,000?
Totally we should be wary of a system that protects 400 and kills 1. Thank you for providing the numbers. It helps me show my point more clearly.
If you are driving on a road you encounter cars. Each car is a potential accident risk. You probably encounter a few hundred cars after ten or so miles. Not every car crash kills, but lets just assume they all do to make this simpler. For the stat you propose, you are talking about feeling uncomfortable with an accident per mile of something around the ballpark of ten miles.
Now lets look at the data. The data suggests the actual miles per accident is closer to 6,000,000 miles per accident. This is six orders of magnitude diverged from the number of miles per accident that you imply would make you feel uncomfortable.
Lets try shifting that around to a context people are more familiar with: a one dollar purchase would be a soft drink and a six million dollar purchase would be something like buying a house in the bay area. This is a pretty big difference I think. I feel very differently about buying a soft drink versus buying a house in the Bay Area. If someone told me they felt that buying a house was cheap, then gave a proposed price for the house that was more comparable to the cost of buying a soft drink, I might suspect they should check the dataset to get a better estimate of the housing prices, because it might give them a more reasonable estimate.
So I very strongly feel we should cite the numbers we use. For example, I feel like you should really try and back up the use of the 400 to 1 number so I understand why you feel that is a reasonable number, because I do not feel that it is a reasonable number.
> Also if this was about safety and not just a bunch of dudes who think they are cool because their Tesla can kinda drive itself, why does "FSD" cost $16,000?
Uh, we are a on venture capitalist adjacent forum. You obviously know. But... well, the price of FSD is tuned to ensure the company is profitable despite the expense of creating it as is common in capitalist economies with healthy companies seeking to make a profit in exchange for providing value. It is actually pretty common for high effort value creation, like creation of a self-driving car or the performance of surgery, for the prices to be higher.
If you are advocating against a system that protects 400 people and kills one, you are advocating for killing people.
(Is Autopilot still limited to divided, limited access highways? Those are significantly safer than other roadways.)
No. Was it ever? All you need is a piece of road that has something which appears to be lane lines. The road to my house is usable despite having no actual paint striping because it happens to have a crack that runs fairly straight up one side and was filled with tar. So the camera thinks it's a lane line. Ta-da!
The thing is we often have discussions about this stuff and I'm trying to advocate for citing datasets to more tightly correlate our words with the evidence that our words correspond to. I'm not trying to say this version shouldn't have been recalled for example, but that I think we should be close to evidence.
In the case of auto-pilot, it was the case that people made the same arguments that are now being made against FSD. I think that makes it somewhat relevant to the discussion, because people previously also made the same claims about safety, but now that we have the data, we can see those claims were wrong. I believe these sort of generalizations, though inaccurate, can help us to make more informed decisions, but I'm not really confident in any beliefs that are made at this greater decision from direct data.
So I think anyone who can provide datasets which correspond with FSD performance rather than autopilot performance ought to do so. That would be really great data to reflect on.
The thing I'm worried about is that no data at all is backing the conjectures - which, given that I sometimes see estimates that I calculate to be many orders of magnitude away from data informed estimates - seems to be the case on Hacker News at least some of the time.
They have a set of regression tests they run on new code updates either by feeding in real world data and ensuring the code outputs the expected result, or running the code in simulation.
It does seem worrying that they would miss things like this.
Here’s a talk from Karpathy explaining the system in 2021:
Though I don’t recall if he explains the regression testing in this talk, there’s a few good ones on YouTube.
I used to think that fact was going to delay self-driving cars by a decade or more, because of the potential bad press involved in AI-caused accidents, but then along comes Tesla and enables the damn thing as a beta. I mean...good for them, but I've always wondered if it was going to last.
I've been using it pretty consistently for a few months now (albeit with my foot near the brake at all times). I haven't experienced any of the above. Worst thing I've seen is the car slamming on the breaks on the freeway for...some reason? There was a pile-up in a tunnel caused by exactly that a month or so ago, so I've been careful not to use FSD when I'm being tailgated, or in dense traffic.
There are only like 16 million intersections in the US. Why not test them all?
Tesla does however collect data on edge cases and then train their system to respond correctly. They can for example trail a collection network to identify things that might be obscured stop signs, then have the fleet collect a whole bunch of examples, hand label those samples, and roll this new data in to the training system. This is explicitly how they handle edge cases.
They can also create a new feature or network and roll it out in “shadow mode” where it is running but has no influence on the car, and then they can observe how these systems are behaving in the real world.
The real issue I guess is when they release a new feature without trialing it in shadow mode, or if they have gaps in their testing and validation system.
Yes. An army of Tesla owners perform the QA, in production.
But in all seriousness they do have some small team that validates then it goes to employees.
Formal verification is obviously better if you can do it. But it's still really really difficult, and plenty of software simply can't be formally verified. Even in hardware where the problem is a lot easier we've only recently got the technology to formally verify a lot of things, and plenty of things are still out of reach.
And even if you do formally verify some software it doesn't guarantee it is free of bugs.
Whether it's a neural network inside or not is completely irrelevant. That's why it's called "black box".
You make a black box test on several thousands (sometimes only hundreds) patients, and if patients who received the drug perform better the patients who received the placebo, then the drug is usually accepted for commercialization.
Yet one isolated patient may be subject to several comorbidities, her environment could be weird, she could ingest other drugs (or coffee, OTC vitamins or even pomelo) without having declared it. In a recent past women were not part of clinical trials because being pregnant makes them very "non-standard'.
Secondly, even after the product hits the market the company is still responsible for tracking any possible adverse effects. They have a hotline where a patient or doctor can report it, and every single employee or contractor (including receptionists, cleaning staff, etc.) is taught to report such events through proper internal channels if they accidentaly learn about them.
I don't know where you get that, most clinical trials last 26 weeks, even in phase III.
and about "more thorough than you imagine" no, most are subcontracted to CROs and the way clinical trials are conducted is messy.
Below is story from the POV of a PI.
But similarly many patients complain about the way they are treated in visits and the lack of interest of the nurse/doctor who receive them.
https://milkyeggs.com/biology/why-are-clinical-trials-so-exp...
Clinical trials also have strict ethical oversight and are opt-in. If clinical trials were like Teslas, we'd yeet drugs into mailboxes and see what happened.
Time to stop driving. That is not normal
As for their QA process, in 2018 they had a braking distance problem on the Model 3. They learned of it, implemented a change that alters the safety critical operation of the brakes, then pushed it to production to all Model 3s without doing any rollout testing in less than a week [1]. So, their QA process is probably: compiles, run a few times on the nearby streets (I am pretty sure they do not own a test track as I have never seen a picture of tricked out Teslas doing testing runs at any of their facilities), ship it.
[1] https://www.consumerreports.org/car-safety/tesla-model-3-get...
1. https://finance.yahoo.com/news/upcoming-tesla-software-2020-...
I have a winding road near me with a speed limit of 35 mph, but 15 mph on certain curves as indicated by a speed limit sign. It ignores those speed limit signs and will attempt to make the turns at 35 mph resulting in it wildly swerving into the other lane and around a blind turn with maybe 30 feet of visibility. It has also attempted to do it so poorly that it would have driven across the lane and then over the cliff without immediate intervention.
Unsupported claims by a manufacturer that compulsively lies about the capabilities of their products except when directly called on it are the opposite of compelling evidence.
Teslas definitely read speed limit signs. I've had mine correctly detect and follow speed limits in areas without connectivity or map data. It also follows speed limits on private drives (if there is a sign) and obeys temporary speed limit signs that have been put up in construction zones.
Frankly, attempting to deflect by arguing that it is okay to release a defective safety critical product to unsuspecting consumers just because nobody else is willing to offer a similar product because they have some moral integrity is a stance that makes the executives at Ford presiding over the Pinto look like angels in comparison. All the Ford executives did was cover it up to avoid having to pay to fix it. At least they did not intentionally release a known defective and dangerous product just to recognize some revenue.
Cognitive dissonance at it's finest.
There is even a former Tesla AI engineer that throws objects in front of the car on YouTube, as a demonstration.
The results are not glorious at all :| (trying to find the channel back if someone knows).
And random public tests too: https://www.youtube.com/watch?v=3mnG_Gbxf_w
This is a basic safety auto-braking. Just feels very wrong to even accept it goes into release.
The guy behind this is known to be untrustworthy, and many of the videos don't actually do what he claims. Notably he refused to release the videos that would prove his claims right.
The reality is that Tesla scores high on all the automated breaking test done by government. The driver however can override this, and that is exactly what is being done in this video.
The test done above has been replicated and the car does break in automated driving. And it gives loud warning to the driver even under normal driving operation and does emergency breaking.
Tesla assumes that if the driver hits the accelerator after the warning, the drive wants to accelerate based on the drivers judgment.
This is what this video shows, notably the person that made this video has refused to release prove that the car actually was in Autopilot, refused to provide evidence that the driver didn't hit the accelerator and refused to provide audio from inside the car so it can be verified that there was no warning sound.
In addition to that, the same person also claims things like 'millions of people will die if Autopilot isn't stopped' and that even under absurd assumptions is a legit insane thing to say.
So what is more likely, that Tesla did something incredibly illegal removing and incredibly important safety critical software that is standard in every Tesla OR that a competitor is doing a PR campaign where they deliberately set up a situation to film a video where they can create maximum damage to Tesla and then advocate for their own solution.
https://www.tesla.com/blog/model-y-earns-5-star-safety-ratin...
So the question is who do you believe, Euro NCAP or a competitor who made a sensationalist viral video.
That's what makes it unfinished...
It's never passed the 'drive from new York to LA with nobody touching the controls' test...