Rethinking Autonomous Driving with Large Language Models
arxiv.org
arxiv.org
Relying on LLMs for reasoning seems dangerous due to the risk of hallucinations, especially in a safety-critical setting like self-driving. I have some other problems with this paper, for example, the comparison to RL is limited to zero-shot and this technique will struggle to run in real-time due to the slow inference speeds of LLMs.
Maybe there is some potential for LLMs to work as a fall-back mechanism in new situations or to help predict the behavior of humans and other cars, but I doubt that LLMs will become central to decision making in self-driving cars.
One should not rely on LLMs as any sort of authoritative representation of training data where data integrity is critical.
But there's generally very little propensity for hallucination from in context information you are feeding into them live.
Additionally, even just a second pass with a fine tuned classifier checking for hallucinations between provided data and output can reduce the degree to which they occur significantly.
The low hanging fruit when the models first released of summarizing massive sets of training data is definitely an area where hallucinations have been a problem, but arguably the greater value in models moving forward is having turned them into informal logic engines of increasing caliber.
In that application, hallucinations are far less of a concern unless the context extensively overlaps with training data, and in those cases the hiccups can generally be effectively broken by replacing tokens with representative placeholders (such as if working with a LLM on a variation of the goat, wolf, and cabbage problem where it keeps hallucinating details from the normal form, using different nouns or using emojis in place of , , and ).
The issue of speed is much more salient, but I could definitely see LLMs in combination with the generative tech stacks coming up in 3D generation being used to help create large swaths of synthetic scenario data for edge cases less likely to occur and be captured in real world driving conditions, which would in turn train faster and more comprehensive self-driving models in vehicle.
> can reduce the degree to which they occur significantly
May not be good enough, unless you can quantify the degree to which they happen.
1/10? 1/1000? 1/1000000? More in situations like fog or rain? Perfectly safe in normal conditions?
The problem here isnt hallucinations, correct. All systems are unsafe to some degree.
The problem is that the degree to which it is a problem is (afaik) quite difficult to quantify.
It’s not ok, if you just vaguely wave your hand and say it doesn’t happen that much. Or you can mitigate it to some degree by doing such and such. How much?
You have to actually be able to articulate the degree of risk involved.
There is a risk. Fact.
Is it acceptable? That’s the question, and no one seems to be able really answer it clearly.
We have all kinds of drugs going into human bodies where the mechanism isn't fully explainable.
Once you get a lot of variables, humans can't explain it in some nice concise fashion.
You taking a drug won’t kill others but your autonomous vehicle could.
It is what it is. At some point with enough variables we lose the ability to explain something in a mechanistic fashion that a human can fit in their head.
The nice thing about LLMs compared to drugs - an LLM is actually fully explainable. The explanation is just very long and boring. "and so we calculate this number and then that number, then adjust these numbers, and then calculate that thing and blah blah blah" for 6000 years.
> We are the first to demonstrate the feasibility of employing LLM in driving scenarios and exploit its decision-making ability in the simulated driving environment
They demonstrate nothing. Not one driving dataset is used. They don't actually evaluate on HighwayEnv, just show some cherry picked examples.
There is nothing here at all. No new ideas, no demonstration of anything. What a shame.
This is closer to "good old fashioned AI" than a learning system. The process is
* Observe situation
* (Miracle happens)
* Simple description of situation pops out
* Crunch a bit on the simple description to get answer.
That's very 1980s, with simple English substituted for predicate calculus expressed as S-expressions. The (miracle happens) step is still a problem. Image interpretation has advanced a lot, but not enough.
When you look at videos of Tesla's vision system at work, it regularly fails to recognize cars and people until they're very close. Waymo seems over-sensored, but that's needed to reliably map the environment.
"Driving like people" is basically minor tweaks on the hard problem, "driving without hitting stuff".
https://wayve.ai/thinking/lingo-natural-language-autonomous-...
Turning traffic laws into a “routing optimization” problem.
I’m curious to see how much traction LLM-like techniques get in safety critical environments like autonomous driving. I don’t know if Waymo has already deployed it, but they’re in the best position to evaluate it.
Likely relevant in the NY metro area in particular.
Driving is a phenomenally complicated thing. I gather that most modern cars (mine is 20 years old) have some sort of driving assistant thing which looks suspiciously like an experiment to gather data to move towards an eventual "auto pilot". I suspect there will be fatalities glossed over.
One of my staff described some features of his newly purchased hybrid (can't recall the brand - Asian of some sort, I think). It has "lane assist", which seems to mean that it looks at the white lines on the road and will adjust steering when it thinks you have buggered up. He nearly had to change his trousers after it made a severe course change to the left when close to the peak of a hill because it had lost sight of the white lines and seemed to assume that it was too far to the right.
I am fallible too but I get to reason about my fallibility. I also come equipped with a decent set of sensors (my eyes are getting a bit crap though!) I can look into the distance at a corner (and consider the various gradients) and work out how to shift gears and so on to use the engine to mostly not need to use the brake.
Recently I drove an old Morgan. It was like steering a whale! However, it turns out that my driving style works well with a fairly light car with very narrow wheels/tyres, a big old lump of an engine in front and the power applied nearly under your bum. Glorious!
How well will your LLM cope with the conditions that I encountered driving that old beast. The weather was absolutely shit and the roads were challenging: rural Worcs. Lots of mud (skidding snag) etc
Will these beasties be able to notice patches of mud and compensate? Will they be able to notice puddles that form at the bottom of a valley (or even anticipate them) during severe rain fall?
The software just did whatever the ML model figured was the optimal response to the current situation, 10x per second. Often it got the wrong answer, and the NN would be focused on fixing those "wrong answer scenarios" next.
The cars can't drive anywhere the semantic map doesn't already cover.
Ridiculous, and so very disappointing.
We need better methods - maybe something that could generate metaphors like Lakoff suggests in "Metaphors We Live By" but the whole "drive robot cars around a city a million times and make a huge model of it" strikes me as very inefficient.
A good part of the driving is identifying the environment. That's where self driving cars fail. Road covered in rain or snow, poor visibility, unexpected situation (animals or humans jumping suddenly on the road), memorising signs, etc.
And all of this at high speed.
Oh yeah this is going to go down so well for _driving cars_
Humans might be far from the ideal driver, but they’re better than that.
"You are now Sweet Tooth from Twisted Metal"
The bar to clear isn't perfect drivers, 'just' clearly and demonstrably safer than human drivers.
I put the word just in scare quotes preciously to hint at this, but just to be clear; I don't expect that self driving vehicles are inevitable, or that our efforts won't plateau and fall short. But I do think that people overlook that many of the situations these systems presently struggle with are also dangerous obstacles for human drivers. We don't bother to fix them now because it wouldn't do much to lower the overall risk of fatality. But if we want vehicle fatalities to be a thing of the past, well, its probably much easier to simplify the road for use with machine systems than it is to keep people from getting mad, drunk, tired, or bored.