This announcement comes after a 4 month sabbatical where Karpathy said he wanted to take some time off to “sharpen my technical edge,” which makes it sound like this is the result of frustration with the technical approach instead of burnout.
This announcement comes after a 4 month sabbatical where Karpathy said he wanted to take some time off to “sharpen my technical edge,” which makes it sound like this is the result of frustration with the technical approach instead of burnout.
The fundamental flaws are in the decision-making being based upon 10-30 second feature memory, ignoring features outright, and only depending on visible road features instead of persisted map data.
For instance, near my house there's an intersection where it will try to use a turn-only lane with a red arrow when it's trying to go straight thru a light. 100% of the time. Even if I'm in the correct lane with no traffic around.
That's because the turn arrow on the ground is worn off. It is not a perception problem, it is
a) ignoring the obvious red turn arrow signal's significance for lane selection deliberately b) makes no attempt to persist or consult map data for 'which lanes go where' c) it completely disregards painted lines on the ground in the "no drive here" striping.
Also one block from my house, FSD will stay still indefinitely waiting for trash cans (displayed as trash cans) to clear the leftmost lane so that it will turn left.
None of the failures I encounter are due to lack of perception.
Tesla has been collecting thousands of dollars, each, from car buyers and utterly failing to deliver what it represented, and keeping the money year after year. Would be let GM, Toyota, or Audi do this? Where is the criminal prosecution? Where are the refunds?
I would not, personally, be so dissuaded.
I bought it knowing it was in beta and with no timeline for production release. They haven't represented it as anything more than that.
Being along for the journey is a big part of why I decided to pay.
Is not
"literally said in a public conference keynote that robotaxis would be running in 2019"
Here is a longer video of that announcement: https://www.youtube.com/watch?v=0GnH_C6NrOM
I think that's optimistic. Ten plus years.
There's too many exceptions. In my home town, an intersection. Imagine, if you will:
Traveling westbound there's four straight ahead lanes. There's a traffic light for each straight lane. The traffic lights will alternate "left two straight, green; right two straight red" and then "left two straight, red; right two, green". It does this because there's a tunnel and a roundabout right there. I guarantee that FSD will choke on this.
We also have a road with a reversible lane in the middle. Thankfully it recognizes the red X sign for this. Unfortunately, if it's driving for 11 seconds without seeing anothee, it will suddenly decide to merge into the incoming lane to make an upcoming left...
I think in Japan or China this approach could actually work, but there is zero hope for it to work generally in the US.
You aren’t beta testing complex automation system that’s operating on public roads. To do any meaningful testing you should have defined operation domain, specific behaviors to test, direct line to the engineering team to report issues, etc, etc.
I'm thinking of being an unwitting test crash dummy on public roads and sidewalks for "Full Self Driving". They just slap the binding legal agreements and waivers on their drivers, but it is tantamount to fraud. I'd rather this didn't exist on public roads in such numbers, and that people don't have to pay to be free Tesla drivers, beta testers, and data points instead of being considered people.
It's not acceptable to have it fail for limited number of users. As my kid who may get killed by the failure didn't sign up for it.
How would you feel if this was a standard for aircrafts? "Let's just push this change to a small portion of the fleet, and boy, if that plane crashes, it's soooo much better if all our planes would crash. High five, where's my bonus?!".
still a very hard problem.
but what it requires is a consistent world model. and if it detects that the current stretch of road is incomprehensible for its model it should safely disengage.
of course Tesla opted for the cowboy version of this, and turned the confidence of the model up .. without having a good model :|
3 years was my minimum, assuming they scrap their current model and begin working on one with real geo-spatial awareness now.
They have the data for it already.
To over-essentialize the problem with their approach, you simply cannot safely drive in Atlanta as a person who has never seen the streets. Humans are no exception. You have to know how each road works, and remember that for next time.
I vividly remember a conversation I had with an acquaintance who had recently taken an engineering position with an AV company that occurred circa late 2018. He claimed that they were at most 1 year away. In fact, his exact words were something along the lines of "They're already here. It's just a few edge cases to work out and some regulatory hurdles to overcome."
The reality is that the first 80% of the problem had been solved quickly and significant progress had been made at the time on the next 10%. The end was in sight. Unfortunately, that next 10% ended up taking as long as the first 80% to solve, and the final 10% will likely take decades if it's even possible.
Elon has been over-promising(i.e. flat out lying) about self-driving every year since.. 2014(there's a youtube video compilation of it)?
It seems like his strategy is to just come up with increasingly grandiose promises every year when he fails to deliver on his past promises. He's trapped in his swirling vortex of bullshit. Very worrying to see Karpathy leaving...
Elon in 2026: “by 2028 we’ll have FTL drives”
Elon in 2028: “time machine!”
Well, to be fair, he only has to hit _that_ goal - at which point he can go back in time at his leisure and fix all the others. And he could hit even the time machine goal as late as he wants, and it won't matter.
So I think the 'other civilization' thing only works if they're already approximately here with a running machine.
Wanna start a trading venture with me? :-)
To be fair he'd only have to hit the FTL goal and then FTL can be used as a time machine.
We might well have one - back to the stone age based on the potential to escalate ongoing conflicts...
Steve Jobs was also known for his “reality distortion field”, so maybe it comes with the territory.
That said, FSD progress seems asymptotic and the Optimus thing always seemed like bullshit.
The only ones I met who thought so was managers and software developers and ML-people, i.e. people who has never in their life seen the level of effort it actually takes to get something qualified for RTCA or other safety standards.
And you were right, but most people aren't like you, and people like you don't drive the discussion on what trends are coming. The world needs more folks with your background. Generally no one wants to listen to folks like you because you come off as a Debbie Downer despite being correct most of the time.
I think Mercedes has a reasonably sensible game-plan as far as I know. They currently have Level-3 on some highways in Germany which is still freaking hard to achieve, but still a narrow enough scope that you might be able to pull it off.
Time will tell, my bet is its still 10+ years off.
The guy was critized for being old school and not having adapted to the latest tech (including by elon musk when he dumped mobileeye), but in fact he was just plain right.
i think the tech community owes him an apology.
But his fundamental opposition to ML for things related to safety was, IMHO absolutely correct.
Let's face it Neuralink isn't going anywhere too.
Tesla wasn't started by Musk and relies heavily on Panasonic battery developments and SpaceX and Starlink is good old fashioned ex-NASA and MIT talent.
What does Musk actually do? Annoy the SEC, go on Joe Rogan to smoke weed and have more girlfriends per hour than Gengis Khan?
Seriously? Either you're trolling or you have some irrational hatred of the guy.
Lets remember the only reason he's around today is because some very rich people and NASA bailed him out. SpaceX left to just Musk's ability, ego and resources died in 2008.
He relies so much on mysticism for parting people from their money, eventually some people get fed up with the 'mysterious future and repeated failure' business model.
Expectations around his personal abilities should be lowered. His track record is not that great and it's a fact people call trolling. Sorry, not sorry.
The abilities that made spacex and tesla were hundreds of highly skilled engineers. Elon just raised money and is the face brand for japanese battery tech.
He's not worth worshipping and his incompetence is covered by idolatry.
“Just raised money” is patently false. He’s known for being extremely hands on with both Tesla and SpaceX.
I maintain you’re either trolling or ignorant.
1) the beams are not full, so many users are getting unrealistic service since they're not doing traffic shaping in those areas yet.
2) the areas that are full are seeing significantly worse speeds and latencies. see Reddit.
3) Elon said <20 ms pings will be normal. they're averaging 30-40ms in full areas.
4) it's not financially viable without many more users, higher costs, and cheaper launches.
They are just going about it better but not trying to selling it.
Any reason why everyone seems to be stuck on this problem?
ML maximalism focused on the narrow problem of 'solving driving' while not recognizing that any task as complex as driving requires probably something closer to general intelligence, and theoretically the field has been impoverished in favor of "throw more graphics cards at everything".
1. Need to understand all standard signage (seems possible with AI).
2. Need to understand all "unstandard" signage (not sure how possible).
3. Need to understand the cop with the thick NY accent yelling at you saying "Can't you see there's been an accident and the road is covered with glass you dufus? Turn the F around."
I can certainly see AI solving the problem of driving in specially designed limited access highways (which could also support normal human drivers), and that alone would be a huge benefit, but I never saw how so many were willing to make the leap to "robotaxis that can drive you anywhere in the city."
5. Need to understand that occluded objects have not vanished from the universe never to be seen again
What you describe is the classic west coast non-confrontationalism. You're probably annoying enough that it's simply not worth it to counter anything you say because you just dive into a petty and self-aggrandizing argumentative mode.
Brash and confused conservatives think everyone agrees with them, but in reality people don't want to get caught up in absolute bullshit by joining any kind of interaction with them. This is a common pattern by now. You would do well to recognize it.
8. Generally understand the true context of things you're seeing. E.g a 4x8 sheet of plywood in the road, a squashed piece of road kill, a large metal pipe, a stopped vehicle, a baby carriage, a dishwasher, a tumbling large slab of styrofoam (all things I've encountered in the last year).
9. Is there a stopped firetruck in front of me (I list this because of the ludicrous fact that a Tesla plowed full-speed into such an obstacle).
- driving across an unmarked grassy mound to park a car at a store in their designated area.
- paying a fee with coins to enter and exit a toll road.
- stopping to move around roadworks based solely on hand signals from one of the workers.
These aren’t even scratching the surface in terms of edge cases that could be encountered regularly.
People point at statements like "it just needs to drive better than the average driver" and "well lots of idiots drive".
On the flip side, lots of intelligent people with other skills do not have the confidence & comfort to drive outside of their towns local roads, or at all.
There are problem sets for ML like quant trading or phone album image tagging where you just need to be right 51% of the time to make money / be novel & useful.
We've seen plenty of examples showing ML is not appropriate for life&death problems like being the primary reviewer of diagnostic imaging for cancer screening.
The bar for ML to drive a car is much much higher, and I'm not even sure we have the right sensor suites feeding into the models, not to mention the compute capacity to operate at low enough latency.
I thought they had real self-driving taxis in Pheonix that you can order? Real ones, with no safety driver.
That definitely sounds "better", even if it is heavily geo-fenced.
Depending on the route, you could probably even do it with comma.ai hardware.
When I think of FSD, I think any route under any condition.
They absolutely cannot. They won't even try, they require a driver to be there to be ready to take over with no notice. Their software also makes so many basic mistakes that even what they're allowing it to do is dangerously reckless.
The Tesla AI Day[0] surprised me as it showed they only had a simple architecture for a very long time, simply feeding barely processed camera pixels to a DNN and hoping for the best with little more than supervised learning off human feeds. Their big claim to glory was that they rearchitectured it to produce a 2D map of the environment… which I thought they had years ago, and is still a far cry from the 3D modeling that is needed.
After all, sure, we humans only take two video feeds as input… But we can appreciate from it the position, intent, and trajectory of a wealth of elements around us, with reasonable probability estimates for multiple possibilities, sometimes pertaining to things that are invisible, such as kids crossing from nowhere when near a school.
Cruise also seems to have better tech; they had a barely-watched 2h30 description of their systems[1] which shows they do create a richer environment, evaluate the routing of many objects, and train their systems on a very realistic simulation, not just supervised training, which means it can learn from very low-probability events. They have a whole segment on including the probability that unseen cars may travel from perpendicular roads; Tesla’s creeping hit-or-miss are well-documented on Youtube.
No, we use both our perception and proprioception when we drive or walk for that matter. We have two accute visual sensors mounted in a housing with six degrees of freedom and a sensor feedback mechanism.
We also have fairly sensitive motion and momentum sensors all throughout our bodies. Additionally we have audio sensors feeding into our positional model.
All this data is fed into an advanced computing system that can trivially separate out sensor noise from useful signal. The computing system controls the vehicle, navigates, and even performs low priority background tasks like wondering if the front door of the house was locked.
We have dozens of integrated sensor inputs. It's just silly to assume we only use our eyes when driving. They're certainly an important sensor for driving but definitely not the only one.
If you ever review your Tesla Dashcam footage you'll see that you can rarely even make out a license plate. The cameras are not even HD, let alone UHD.
Further the refresh rate is a paltry 36fps.
Drive the car at high noon or night and you realize how poor the dynamic range is with highlights and/or shadows get lost.
Or we could build 1-dimensional roads, which would make the AI's jobs much easier. Like, we could put down two parallel piece of metal, which vehicles could "hook" onto somehow...
I certainly wouldn't argue with you that it isn't ready for prime time and wide distribution, but it is interesting to see their progress in San Francisco, a much different driving problem.
If it takes them 10 years to get to prod in Mesa, two (maybe three?) in SF, maybe they start shrinking that a lot in metros without winters. ¯\_(ツ)_/¯
To solve the self-driving problem we need "smart" A.I., which means we have to approach it with systematic engineering, and the solution will probably involve some combination of better sensors, introspectable neural nets, symbolic A.I., and logical A.I.
Things like traffic signals that actively communicate their status to nearby robot cars (more than just a red lamp that can be occluded by weather, other vehicles, or mud on the camera lens). Or lane markings that are more than just reflective paint, but can be sensed via RF. Rules around temporary construction that dictate the manner of signage and cone placement that the robot cars can understand. The cones might have little transponders in them, I don't know.
But without a massive leap forward in AI capability, our current road system—optimized for human drivers over the past century—is not going to work.
If we can't make the cars just as smart as an alert and capable driver, then maybe we need to meet halfway and make the roads a littler "dumber" (simpler) to accommodate the robots.
Or even cycle? I hear great things.
A part of the reason people wish they had FSD so badly is because they want to be rescued from this fundamental failure of NA-style urban planning that necessities driving, all the time, across both short and long distances.
I heard the other way - people leaving cities for suburbs
It doesn't work like that in most countries; small cities and towns are navigable on foot and public transport.
If you could replace double-decker buses that arrive every 15–30 minutes with self-driving minibuses that arrive every 3–5 minutes, that would be great! (for everyone except the bus drivers who lose their jobs)
Actually, I suspect this makes self-driving a particularly _bad_ solution for city buses; getting into the bus stops takes some manoeuvring, particularly when there are other buses there.
One place that self-driving buses could be interesting (and indeed there are already a couple of systems like this) is on fully/near-fully segregated lines, where they don't have to deal with human-operated traffic. Another would be small towns, but you're looking at full magic level 5 at that point.
I’m envisaging a system of minibuses, either AI-driven or at least dynamically directed by a central control system. People would use a phone app to book journeys; the app would tell them where to get on and where to transfer, and the central control system would optimise the fleet to get everyone where they need to go.
With much smaller buses, and electronic tap-in rather than cash payments, stops should be fast enough that buses can just queue up in order at each stop, hopefully mitigating the parking difficulties you mentioned (modulo breakdowns, medical emergencies, etc). Likewise, if transfers are fast and easy, hopefully they’d be less objectionable to travellers.
Requiring a phone isn’t ideal as it limits accessibility and privacy. There could also be pre-printed tickets, with QR codes that you scan at the bus stop to see the route info.
Why minibuses and not just taxis? I suspect there’s a good balance to be made between efficient road usage (buses) and efficient routing for each traveller (cars).
I’m also envisaging that if this system were to take off, personal cars could be gradually removed from city centres! Again, that makes life easier for the AI vehicles.
This is all pie-in-the-sky stuff, I know; but I do feel like there are ways we could radically improve city transport mostly using existing roads, rather than building new rails or tunnels.
The problem you then face is that any of those could be forged / faked without some kind of way of securely validating the message in some way. You could cause absolute chaos by driving down the road broadcasting false messages. It's a little harder to hack and modify traffic light signals, for example. But we've also seen hackers screw up Tesla cars by sticking stuff on the back of their car to deliberate mislead it based on vision.
Any reason routine crypto methods would not solve this? Seems like one of the easier parts to me.
The real concern would be whether someone can engineer a terrorist level mass scale attack but as long as it requires physical tampering that adds up to a tremendous amount of work. So if the signalling is largely burned into fixed infastructure it eliminates a lot of that or at least sets the bar high enough that its probably more work than various other types of attack that are likely to be just as impactful.
Even without self-driving cars, an "attacker" can go into a theater and yell "fire" and cause a stampede.
They can get a high-viz vest and clipboard, and stand in intersections directing cars to take detours they don't need and holding up traffic.
My point here is that society has a lot of trust baked in. We trust people don't just yell "fire" without reason. Just because it's FSD cars doesn't mean people will start broadcasting the equivalent of "fire" constantly. It's already easy to cause accidents.
The consequences of potential attacks on centrally-orchestrated traffic are a lot more severe. Hack the control node, and you can stop traffic nation-wide. Or cause mass accidents that overwhelm first responders. And they can be executed by anyone, anywhere in the world, for a cost within range of many medium-size corporations (let alone nation states).
I won't comment on the challenges of the approach Tesla et al. are currently taking, but I don't think central control is the panacea commenters in this thread are making it out to be (and I'm personally glad this isn't the route we're pursuing).
It’s like arguing that we can’t possibly build autonomous cars because then someone might turn it into an autonomous bomb.
Keep in mind that solving this is “worth” about 40,000 lives a year in the US - nearly $1 trillion in economic damages a year in life and property.
Bad things can always be done with good tools. As always, you provide layers of protection that make sense and in the end must rely on the underlying fabric of civilization to persevere.
Sounds like a job for...blockchain.
I’ve personally tested creating a fake toad sign and my tesla reads it as a real sign just fine.
I wonder what happens to autopilot on a 70mph freeway when it encounters a 5mph limit sign…
It immediately reduced speed from 60mph to 25mph .. aggressively.
This falls into a pattern of Tesla autopilot/NoA where it just doesn't seem to have much memory or foresight.
For example the car is driving itself on the highway, it knows it's been on the highway, for 20 minutes. I am not even in the exit lane, it knows what lane I am in. How could it think I am suddenly on the local road below the highway based solely on the GPS pin movement in the span of a second, without having moved to the exit lane and gone down the exit ramp?
For an example of lack of foresight - the car will happily speed towards an obvious semi-distant slowdown right until it needs to aggressively break from 60mph down to 30mph as it approaches following distance of the nearest car. I also find it can get really weird in stop&go traffic, not easing into speed, down to a stop very well as if it has only GO or STOP.
Oh wow. This happens often with a car-mounted GPS (or on a phone) and it's pretty annoying. Sometimes the GPS instructs you to do a U-turn at the next available fork in the road, and it takes a moment to understand what's going on.
But in a self-driving car it's terrifying! And absurd.
Fortunately thats not a lot of people.
When electricity was being rolled out, Westinghouse and Edison didn’t bellyache about not being able to provide electricity to rural areas. They electrified all the cities.
Rural areas will just never be able to pay for modern infrastructure. And… thats ok, its not a lot of people.
Yep, I talk to people working in traffic engineering, and their mindset is always building new road tech and road-side and cloud infra to support autonomous driving. They have no expectation of fully autonomous vehicle without road and infrastructure assistance.
And from historical perspective, the coming of automotive and the replacement of horse and other animal carts, are exactly facilitated by the road transformation; which has been the single largest scale infrastructure in human history.
It makes no sense that an even bigger transformation of the vehicle would require less drastic road transformation.
The very marginal benefits of laying road vs track more or less disappear when automation in play.
Is is really worth maintaining the ability to go off road/track when you’re not even driving anymore?
This idea is not new and it mostly applies to freight convoys but I think it also has merit for ad hoc passenger car convoys on long highway trips.
I think the idea is to get rid of the roads and have the robo vehicles travel instead on tracks. That increases fuel efficiency and bypasses a lot of AI challenges.
It’s the low-key case that Elon will be remembered poorly (like Robert Moses’ rapidly degrading legacy) for
1) Having the wrong vision for EVs (but successfully executing non it anyway)
2) Making space travel cheap (thereby increasing the amount of carbon energy dedicated to it) without really improving an average human’s quality of life
Are our roads really? Most in cities over a certain age are just haphazard relics of times gone by, and don't get me started on "stroads" which are good for nobody
This is a thing in Europe, and even some US cities - my Audi has traffic sign recognition and when at a compatible intersection knows what the light is at (by radio, not by light), and how long until it changes (will show a countdown in seconds til the next light change).
what's missing is combining this kind of "human concept relations" model (language, rules, minimal reasoning, text encoded human preferences) with perception, and safety (which means that the model should know that if other cars are driving just fine in front then it's unlikely that the road is on fire, or that the low certainty crack in the road is okay if two other cars already went over it unimpeded, if the road marks and the signs are inconsistent, but other vehicles have formed a slow but consistent pattern of traffic then that's the local ruleset, and so on)
it's still a very hard problem. and the required amount of compute is still bonkers, the required amount of data and training is still absolutely huge, and the whole problem of safely disengaging, handling the asleep/drunk passengers (likely target audience after all)... are all hard problems too :)
This aspect of FSD has always fascinated me and I'm a little surprised it doesn't get more discussion. Meeting halfway. At what point could/would/should FSD influence the environment around it?
For example - a poorly painted road sign*. Tesla/Waymo could say "We cannot support L5 FSD on this road until you fix this sign." If it meant a step forward in autonomy, Tesla/Waymo could even offer to share the cost of that improvement!
There are a million reasons why implementation of that would be problematic. Costs and incentives would be all over the place. But I am more interested in the framing: The machines are the ones that need to adapt. Which is essentially hoping for continued hardware improvements or a spaghetti mess of if/else statements. ie "do this weird thing if you see this other weird thing in front of you". Can we get rid of the weird thing and avoid the engineering challenge altogether?
* Yes, this is an overly simple example. Some environment changes could be so large that they would require a full redesign of a city/buildings/traffic patterns. But surely there are classes of improvements where some are easier than others.
Edit: I have a whole mental model for other drivers and different approaches for them. Someone driving like a grandma? Pass when available. Nervous/erratic/lost driver? Keep extra distance then pass as soon as possible. Aggressive driver? Relax, give some space and let them get ahead. And so on. I get that stereotyping is bad but ignoring the subtle signals other drivers give off seems like it would be myopic. An AI that doesn't anticipate what others will do on the road will always be reactive rather than proactive.
They’ve been absorbing years of data classifying traffic and things like that.
I maintain the position that if self-driving cars handle the most common situations as well as average human beings, there won't be a strong drive to make cars that drive significantly better than humans. Cruise and Waymo are collecting a limited version of that dataset right now. It's unclear how many deaths because some details are being kept under wraps.
We can't build a robot which can walk down a sidewalk without running into people either. The sensor tech and mapping fidelity are red herrings. People drive well because only people are good at predicting human behavior.
Alternatively, an autonomous vehicle operator in a homogenous network full of other autonomous operators has capabilities and characteristics that greatly simplify failure modes. Maybe even majority autonomous, partially heterogeneous? You can literally slow or stop the whole show to deal with a catastrophic event. It’s still “social” but probably much reduced from the scenario where you’ve got the full scope of human expressivity behind the wheel.
The REAL problem is how do we take our roads to the crossover point where those simplified network features become accessible.
I not only assume it, I say it out loud: “[fully autonomous, …] maybe even majority homogenous, partly heterogenous?”
> “When our roads are not used exclusively by motor vehicles to begin with.”
I’m not following what you’re saying here in the context of the earlier clause. If you mean not used exclusively by autonomous vehicles, yes, and that’s why I’m pointing out the provisional aspect.
But I’ve thought the same, too. If every car on the road is robot-controlled then it changes the problem. Modulo failures, discrete algorithms should behave predictably towards each other, like the unix API philosophy.
It seems hard to get there, though. Even today it’s a PITA to maintain API boundaries in simple libraries, never mind make sure that the new Tesla v12.4 Full Self Driving For Real This Time doesn’t trigger edge cases in Volvo v7.7a Actually Real Self-Driving We Promise.
Can we make software that allows cars to behave as predictably as rail cars, but without the rails? Maybe, but I expect only on limited-access freeways. I’m sure these robot-driven cars will remain incompatible with common road uses cases like pedestrians, cyclists, and children chasing balls.
https://np.reddit.com/r/CatastrophicFailure/comments/43juk4/...
Because it's really, really difficult. A lot of AI-ish stuff pretty rapidly gets to the point where it _looks_ quite impressive, but struggles to make the jump to actual feasibility. Like, there were convincing demos of voice recognition in the mid-90s. You could buy software to transcribe voice on your home computer, and people did. And, now, well, it's better than in the mid-90s certainly, but you wouldn't trust it to write a transcript, not of anything important. Maybe in 2040 we'll have voice recognition that can produce a perfect transcript, and human transcription will be a quaint old-fashioned concept. But I wouldn't like to bet on it, honestly.
And voice recognition is arguably a far, far easier problem.
It's somewhat unfair because we used to simply collect less comprehensive data on performance and therefore know less about our corner cases - but you don't get to live in the future without dealing with the problems of the future.
Is this the fault of software or people generally frequently using bad sound equipment in poor and noisy conditions, such as talking on the phone while driving, on the street, poor connection, wind/rain, etc. ?
Personally, because English isn't my first language, I frequently "fail" to transcribe what is being said and have to ask the other person to repeat themselves. It seems like an AI voice transcriber in 2022 is going to work better than me.
Not sure what you're talking about really.
I took part in a legal deposition where an "AI Transcription Software" was being used. When I received the transcript it had numerous errors, but they were all subtle. More common names were inserted in place of the name that I said, e.g., "Kennedy" instead of "Kemeny". "You have a [something]" was transcribed as "I have a [something]", completely reversing the meaning. And many more errors.
The common thread between the errors was that what was inserted into the transcript would have been the MOST EXPECTED word or phrase, instead of the ACTUAL MORE SURPRISING (surprising in an information-theory way) word or phrase. It's evident that on top of the phoneme recognition layer, this transcription software checked questionable items against tables/graphs of most likely words to occur near the other words it confidently identified in that context. Makes a transcription sound great, but it is WRONG.
The result was that the "AI Transcription" actively destroyed key information and hid that destruction under the guise of a smoothly edited transcription.
Although this surely was not the intent of the system's creators, I cannot think of a better way to make a more evil transcription system.
Tesla is taking a fundamentally more broad and deep approach - working with the fundamental fact that a pair of visual sensors and a compute engine (eyes & brain) can successfully figure out driving in strange areas in real time, ergo, it should be possible without a map/model or lidar. Once they get it solved, it will be solved once and for all. Bigger gamble, bigger payoff. Equipping the car with dozens of eyes is the easy part. The question is whether enough compute power can be brought to bear on solving the recognition problems, and the edge cases. They have obvious issues with failing to recognize large objects like trucks in unexpected orientations, left turns etc. Using millions of miles of live human driver data as a training set is great, except that the average driver is really bad, so it's entirely polluted with bad examples, ESPECIALLY around the edge cases that get people killed. There, examples from professionally trained drivers, who really understand the physics and limits of the car, adhesion, traffic dynamics, etc, are what you want to train on, but that isn't what they have. It is also possible that even if the set of training data would actually be sufficient, the big question will kill them - perhaps the solution requires orders of magnitude more compute power to approach human performance, and they just don't have the hardware to simulate human compute power. So, have they just hit the limits of what their compute power can do?
I think Tesla's approach is fundamentally the way to go, as it is a general solution, compared to everyone else's limited map/model approach.
But both may require either or both a more specifically programmed higher-level behaviors, and/or something much closer to AGI than exists, something that has actual understanding of the machine-learned objects and relationships, which does not yet exist (if one is known, pleas correct me - I'd love to know about it).
"A Cruise autonomous vehicle ("Cruise AV") operating in driverless autonomous mode, was traveling eastbound on Geary Boulevard toward the intersection with Spruce Street. As it approached the intersection, the Cruise AV entered the left hand turn lane, turned the left turn signal on, and initiated a left turn on a green light onto Spruce Street. At the same time, a Toyota Prius traveling westbound in the rightmost bus and turn lane of Geary Boulevard approached the intersection in the right turn lane. The Toyota Prius was traveling approximately 40 mph in a 25 mph speed zone. The Cruise AV came to a stop before fully completing its turn onto Spruce Street due to the oncoming Toyota Prius, and the Toyota Prius entered the intersection traveling straight from the turn lane instead of turning. Shortly thereafter, the Toyota Prius made contact with the rear passenger side of the Cruise AV. The impact caused damage to the right rear door, panel, and wheel of the Cruise AV. Police and Emergency Medical Services were called to the scene, and a police report was filed. The Cruise AV was towed from the scene. Occupants of both vehicles received medical treatment for allegedly minor injuries."
Now, this shows the strengths and weaknesses of the system. The Cruise vehicle was making a left turn from Geary onto Spruce. Eastbound Geary at this point has a dedicated left turn lane cut out of a grass median, two through lanes, a right turn bus/taxi lane, and a bus stop lane. It detected cross traffic that shouldn't have been in that lane and was going too fast. So it stopped, and was hit.
It did not take evasive action, which might have worked. Or it might have made the situation worse. By not doing so, it did the legally correct thing. The other driver will be blamed for this. But it may not have done the thing most likely to avoid an accident. This is the real version of the trolley problem.
[1] https://www.dmv.ca.gov/portal/vehicle-industry-services/auto...
[2] https://earth.google.com/web/@37.78169591,-122.45337171
[3] https://patch.com/california/san-francisco/speed-limit-lower...
Just to be clear, this statement only applies to AVs, right?
For all the woe and gloom in the news reporting, Google (and Cruise)'s rollouts have been more or less what I expected: no enormous accidents that were clearly caused by a computer, but instead, a small number of small accidents usually due to the human driver of another vehicle doing something wrong. That seems to lead towards greater acceptance of self-driving cars and confidence that they are roughly as good as an attentive newbie.
The next big situation, I think, will be some really large-scale pileup with massive damages and deaths, and a press cycle where the self-driving car gets blamed. But the self-driving car collected a forensic quality audit log, which of course will aid the police in determining which human caused the accident.
Humans tend not to do that, and, as a result, some fraction of the time they get T-boned.
AI should absolutely mimic the behavior of real (good) drivers.
Although that's not what you're describing here, another problem for AI could result from it knowing more than an average driver; for example, if a high-mounted LIDAR were able to see around corners and let the car decide it's "safe" to do a turn that no human would attempt for lack of visibility, that could cause problems.
(Also, it's surprising that an autonomous car doesn't detect that another car is following it too closely, and slows down appropriately in anticipation. How is this not taken into account.)
An autonomous vehicle has to protect its occupants at the expense of everything else (or, at the very least, appear to do so in a convincing manner), because otherwise no one will step inside.
(At the very least, if a machine is going to sacrifice me or my family to save a third party, I need to know what hierarchy it is following, and how it was decided and by whom.)
But what this incident seems to illustrate is that it's difficult for a self-driving car to share the road with human drivers, and behave like a human -- meaning, allowing human drivers anticipate what it will do.
I drive a motorcycle and a bike in Paris; the reason I'm still alive is because after so many years of this I can tell what all the other cars will do at all times, before they know it themselves.
But an autonomous vehicle that would behave so differently from a human driver as to be unpredictable, would be terrifying.
Even the simple automatic emergency braking features in my model 3 result in some dangerously unpredictable behaviour at times. I have them set as off and insensitive as possible and they still do some awful stuff from time to time. I was on a road trip in northern Ontario this week passing a service vehicle moving very slowly along the shoulder on the right. I was doing 105 km/hr in the right lane of a 3 lane road: 2 lanes in my direction and 1 opposing. The 2nd lane switches directions every 5-10 km to serve as a passing zone. Posted speed limit is 80 km/hr, but prevailing speeds on this road are 90-110 and you’d never be ticketed for anything under 110 in this region. There was a truck gaining on me coming up on the left, so I signalled left and partially moved over to give the service vehicle some room, but didn’t fully take the left lane to let the faster truck know I would let them through as they came past. Very common pattern on this type of road and circumstance. The AEB slammed on the brakes as I came level with the service vehicle despite the fact there was plenty of space to complete the pass. I wasn’t expecting my car to slow down let alone apply emergency braking force, so in my surprise I nearly collided with the service vehicle. I have no idea what the other drivers thought, but neither of them could have possibly been predicting or expecting me to slam on the brakes. I was really upset with the car since it took a highly dangerous action in an otherwise perfectly safe and common situation. And that was the automatic stuff that can’t be turned off; I don’t let autopilot drive ever because it is the worst type of driver: unpredictable. The choices it made in the extremely short time I tested it about lane placement, follow distance, defensive driving (and the complete lack thereof), and general behavioural clues provided to other drivers were genuinely terrifying.
It performed the only sane solution to the problem - stopping. You can't possibly predict what the most likely thing is to avoid an accident, because there is a multitude of factors beyond just 2 cars moving towards each other. Are there pedestrians present? Other parked cars? Storefronts with customers inside? How close to the sidewalk? Trees? Construction/Debris?
Once we get AI working better than 99.9% of humans on roads with speed limits of up to 30 mph - we can expand it to faster roads and introduce advanced behaviors. But for now, stopping in uncertainty is the best available option.
Not only there are other factors, but you also need to predict what the other human driver will do. As stated in a sibling comment this is a huge part of safe driving.
> Once we get AI working better than 99.9% of humans on roads
But what you stated above require AGI, so what's the plan to get there? It's even more blurry than commercial fusion at that point.
Because they're all trying visual- or line-of-sight methods only, I call this the "robo-human" fallacy in ML: trying to automate the processes that humans undergo so that you eventually have a drop-in replacement for a human. But that is a myopic and unimaginative approach because you could be re-assessing the system itself and eliminating inefficiencies that lead to poor performance.
In the autonomous vehicles space, there is massive potential for self-organizing swarm algorithms to control pelotons of cars, rather than individual cars with no intrinsic sense of the general flow of traffic. You wouldn't need a top-down "commander" style architecture, it could be designed so that cars only talk to their immediate neighbors and emergent patterns keep traffic flowing smooth and fast.
I have always been skeptical of the attempts to reduce the amount of information about the road that a car receives. (Moving from stereoscopic to monocular vision to save the cost of one camera seems just stupid.) But people who dream of "smart cities" really seem to see little more than The Jetsons in their mind, and it limits the scope of research to our detriment.
How is AI supposed to confidently distinguish a real stop sign from someone/something holding up a picture of a stop sign?
Yes, this is a weird edge case, but I think it gets at the core issue being that it takes way more sophistication to release this tech into the wild then ppl would like to admit.
They may want to think about that strategy soon. Model 3 is starting to seem dated (not to mention Model S, which is ten years old). There are very competitive alternatives on the market now that have strengths where Tesla is weak, and which are not especially weak in the areas Tesla is strong.
https://en.wikipedia.org/wiki/Bj%C3%B8rn_Nyland_(YouTuber)#W...
Range is okay, but Tesla overestimates. Mine lost about 1.2 miles of range for every mile driven, and as a practical matter it was more of a 200 mile car than a 300 mile car, unless you really wrung it out. There are a bunch of 300+ mile EVs either on the roads now or releasing in the next few months. The only Tesla that has a range worth bragging about is the 400 mile version of the Model S. I do look forward (hopefully!) to regular cars having that kind of range, instead of just the top end ones.
I personally think Tesla is playing with fire leaving the quality so low. Well over half of all Model 3s have to return for repair within the first month. That's terrible. Recall how long it's taken other domestic manufacturers to regain any kind of reputation for quality, and even then many people will never believe they make good cars. For what a Tesla costs, people have high expectations. They're enjoying fad status right now, and that's great, but there is no shortage of Tesla owners already who've sworn off ever buying another one.
There are certainly some questions that are hard to answer. Maybe the Tesla will continue to hold value. But given that I see them everywhere now, they're feeling more like a commodity every day. I'm not sure how it will play out, the market is dynamic and recent events have been very disruptive. All EVs are selling out months ahead of production at this point.
They have terrible manufacturing quality. The screens melt in Arizona heat. Maybe the ride is cool and feels good, but the car itself is not incredible.
That's a very old talking point about the original Model S screens from ~2012. Do you have any recent justification for this claim?
This is not to helpfully make sure the driver gets into a nice cool car when they finish their errand. It’s because the non-automotive grade screen isn’t rated or tested above 40C and were failing in heat. Rather than upgrade the part and do a recall Tesla opted to just use battery to keep the cabin below 40-45C.
I love EVs, but that’s some pretty dodgy behaviour.
This is not really an option in an ICE car because it would require running the engine. In an EV, it's easy.
They have a big head start, but other car companies are now investing much more in battery tech etc and will quickly catch up. Not to mention Tesla's have terrible build quality, they have a lot of shady business practices like overcounting sales, reusing sold parts etc which came out in the recent leak.
Thing like the 4860 battery which were so hyped turn out to be not that much better. FSD is years away. Stop selling vaporware.
What they need to focus on is things they innovated on like OTA updates, integrated systems, no dealerships etc.
How so? They're not selling robotaxis or building factories to build them
> Tesla has repeatedly promised FSD is right around the corner
Which means it's years away and/or "FSD" means "automatic cruise control and lane keep assist" or whatever standard feature from auto manufacturers they've renamed
Because they chose to back themselves into that corner. Musk says that Tesla is worth nothing without full self-driving. Certainly it's the only thing left to justify the stock price:
https://electrek.co/2022/06/15/elon-musk-solving-self-drivin...
> Which means it's years away and/or "FSD" means "automatic cruise control and lane keep assist"
Well, more precisely it means Musk has been lying about it for nine years straight:
https://jalopnik.com/elon-musk-promises-full-self-driving-ne...
The lies have been profitable so far. People have bought into the false promises. Perhaps they'll start demanding refunds for the full self-driving they paid for that has still not been delivered.
Solar
And I also think they could do some clever stuff with home HVAC, possibly using waste heat from crypto miners as the H.
Lastly, afaik they went camera-only in their cheap cars (3/Y) and still use fancy stuff in the S/X.
They gave fsd customers new computers once. What’s to stop them from going back to vision+lidar or whatever once the parts are available and retrofitting as needed?
The whole ‘vision-only is better’ gag seemed like an obvious ploy to keep being able to ship cars from the beginning of supply chain problems.
And yeah, Elon does not appear to be a good person.
> Tesla has just confirmed that it has removed radar from the Model S and Model X as of mid-February 2022, moving its entire lineup to what it calls ‘Tesla Vision,’ which is an array of cameras that Tesla says negates the need for radar. The manufacturer did the same for the Model 3 and Model Y in May of last year and even though this prompted some questions from the IIHS, the institute is now fine with it after testing.
That doesn't sound... particularly clever. Crypto miners are essentially disposable, with a useful economic lifetime of a couple of years, typically. And they are 100% thermally efficient. Which sounds quite good, until you consider that an air source heat pump is typically around 300% thermally efficient, rising to up to 500% for ground source.
So, once your crypto miner is no longer really making anything (a couple of years), you're left with a really inefficient heating system.
That's exactly what I would expect someone burning out to say. You feel the burnout so you need time to get over it and feel 100% (regain your technical edge). You're still burnt out after 4 months, so you don't come back.
Frustration with the technical approach can also cause burnout.
There is usually a hierarchy of sensors, mainly for redundancy. Example: Bumper sensory at the wheel base, sonar / Lidar at the mid, and a camera at the top for advanced sensing.
For the sake of cost cutting Tesla has done away with their radar sensors at the front of the vehicle. It would be a substantial cost overhead, but have very real repercussions when it comes to safety, while also providing a "ground truth" to what at least the front facing cameras are seeing.
I don't think Lidar is a practical sensor for them to adopt, because it is quite bulky and has limited viewing angles, but I would expect them to have adopted some novel, lower cost radar solution.
Apart from the lower cost of the camera, I think Elon's rationale for having a camera only FSD is not valid, has made the problem needlessly complex and unsafe. He believes since we have eyes, and we can drive a car, then it should be sufficient to drive the car, but we only use eyes because these are the sensors we were born with, it is the best we have. In my mind, Elon's approach is like looking at a horse, and saying to yourself, that you want to build a car based on a horse, where instead of wheels, you have four mechanical legs, and those mechanical legs are limited is so many ways, but they should still at least "work", but there is no reason to limit locomotion in that way. The same with the vision system on a FSD, the whole spectrum of light is available, with any number of configurations, providing data at rates and with precision far beyond what a camera system can do.
My background is in physics, but I find myself having a growing appreciate for the vision-only stack. It's really challenging building a formal understanding of the world that is robust to outliers that are so numerous as navigating in an urban environment. With vision, you have multiple kinds of information that are highly correlated (colour, spatial distribution, depth, etc) that are self-consistent. Whereas, fusing radar with vision, where object responses to radar are highly geometry & material dependent, is a much harder task.
I'm really not an expert, so this reads more as an opinion than an experienced view, but I can see the merits in doubling down on vision.
And Tesla cars have more than one camera on them. The front-facing camera is actually an array of 3 cameras (the two farthest ones are at about human eyes distance), but they're also equipped with forward and rearward looking side cameras, and back cameras.
I think Tesla underestimated how hard vision-only FSD is, but having a single camera (they don't) is not the reason.
Also, I never said they have one camera. Multi camera != multi view.
LOL no, he was jumping ship already.
BTW, Andrej, if you're reading this, it is not just excellent it is beyond excellent. I do a lot of tinkering with transformers and other models lately, and base them all on minGPT. My fork is now growing into a kind of monorepo for deep learning experimentation, though lately it started looking like a repo of Theseus, and the boat is not as simple anymore :)
Well, I'm not sure that anyone's tech stack is capable of solving it; the live examples of robotaxis are, well, not something you'd bet your company on (and generally their creators are _not_ betting their companies on them). There was, I think, a decade ago the idea that fully self-driving cars were a near-term inevitability. That's fading, now.
Both clauses seem wrong.
Essentially an admission of failure for the self-driving car industry.
How so?
If humans can master driving with 2 eyes looking forward, why would a car with plenty of cameras in all directions not have sufficient sensory input to master it?
The problem is the software, not the sensors.
If birds had engines on their wings, they would probably fly more like airplanes.
But cameras are a poor substitute for human vision, because they can't move or pivot or refocus much.
you're using elon's own argument btw, are you repeating that knowingly
I think Karpathy realized (probably way back) that cheap sensors + no HD maps + their (reckless) public testing feedback loop doesn't advance towards L5 self driving and is bailing out. Karpathy has always backed Elon Musk whenever he talks about their technical approach, so it can't be frustration with the approach all of a sudden.
Tesla filed with the FCC in May to get authorization for a new radar system.