Building the Software 2.0 Stack by Andrej Karpathy [video]
figure-eight.com
figure-eight.com
If this guy was working on adtech, that would be fine. That's very error-tolerant. But this guy is working on automatic driving.
The basic mindset here is to run image classifiers to classify the objects in an image, then use the classifier output to decide what to do. There's no geometric analysis. That's scary. Classifiers just aren't that good. See the earlier article today about adversarial attacks on classifiers. Classifiers pick obscure details of images and use them to make decisions. Nobody seems to know yet how to prevent that. This problem shrinks with larger data sets, where hopefully the irrelevant details cancel out as noise, but, as the speaker points out, that breaks down when you have few training cases of certain situations.
The Google/Waymo approach is to get a point cloud with LIDAR and radar, profile the terrain and obstacles, and figure out where it's physically possible to go. That's geometry based. In parallel, a classifier system is trying to tag objects in the scene, which feeds into a system which tries to predict what other road users are going to do.
With that approach, a classifier result of "not identified" is fine. The system will detect and avoid it, or stop for it, and make conservative assumptions about its expected behavior. Chris Urmson, in his SXSW talk, showed video of a woman in a powered wheelchair chasing a turkey with a broom. This was not identified by the classifier, but it was clearly an obstruction, so the vehicle stopped for it. That's essential here. It has to do something safe with unidentified or mis-identified objects.
At Tesla, Musk insisted that this could be done with a camera alone because humans can drive on vision alone.[1] So Tesla has people trying to make camera-only driving work. Not very successfully so far.
"November or December of this year (2017), we should be able to go from a parking lot in California to a parking lot in New York, no controls touched at any point during the entire journey." - Musk, in April 2017. This guy is saying what Musk wants to hear.
[1] https://blog.ted.com/what-will-the-future-look-like-elon-mus...
Apparently Musk heard it (or something similar), because Tesla is rolling out the first self-driving features in August: https://news.ycombinator.com/item?id=17282006
Self-driving features? Sorry, there are no some self driving features. You either have full self-driving or you don't. If you don't (which Tesla doesn't), how the hell can you call it self-driving?
In any case, it's quite easy to imagine a plausible meaning "some self-driving features": perhaps these features enable the car to drive itself [without oversight] in some, but not all situations (<i>e.g.</i> on highways, but not in cities).
If you can't detect rain with reliability, you're not ready for full self-driving.
Edit: not to mention state-of-the-art sensor fusion, with atmospheric composition analyzer, thermometer, multiple high resolution vibration sensors, stereo microphones and an advanced vehicle to vehicle signalling system backed by heuristic fallbacks.
One or two fixed position monochrome, low resolution forward facing cameras such as in a Tesla just won't compare.
I'm not saying it can't be done, but the Tesla engineers have a much harder task than if they'd had human level sensors. The Tesla "brain" is more like a human looking out of a tank through a periscope in a tank or something... and the periscope can't even be turned.
I'm pretty sure humans could cope with that. https://www.youtube.com/watch?v=-CITIXlw_T4
(3-camera depth is more reliable than 2-camera. Many of the ambiguous situations for two cameras can be resolved with three. Especially if it's 3 cameras in a triangle, not a line.)
Monocular depth perception w/out motion is a thing too, e.g. [1], however I doubt it is good enough for safety-critical systems like self-driving cars.
Minute point of irony, Musk seems to act more and more like geohotz which he criticized.
There's a great presentation from Gabe Sibley (now with Cruise) on camera based SLAM, or "Mobile Robot Perception for Long Term Perception in Novel Environments" that explains quite well the principles underlying how it works.
Gabe concludes at the end that probabilistic perception and modelling using cameras isn't reliable enough and that Lidar is likely necessary for safety critical systems such as autonomous vehicles:
Agreed it needs more attention, but - for academia - I think it's more of an incentive issue than a stigma issue. E.g. harder to benchmark the performance of two algorithms if they don't operate on the same dataset. Also to be fair, research into things like synthetic data mitigates the problem, just in a different way.
The paper you cited is interesting. Thanks for sharing. Hopefully that spawns more focus into understanding the subtleties of each dataset. IIRC Kaggle also had issues around generalizability, but for different reasons.
Anyways it's still early on... but we're currently building tools to help solve this problem. In particular simplifying the data collection / labeling process for vision systems. Would love to chat further w/ anyone interested in providing feedback. Email is sara@viewpointrobotics.com
I’ve spent time in both academic research and industry.
Research is not supposed to be immediately applicable. The goal is to produce new knowledge - more importantly shared knowledge. Publishing is not a bad measure of that. Additionally, ability to secure grants provides incentive to focus on problems others want solved.
No incentive system is perfect, but I don’t really see how this is any different from any organization. And I don’t think it’s fair to judge an entire discipline by the negative examples.
We still break problems down when solving problems that can apply machine learning. There's no single "drive the car" neural net, rather the task of driving a car has been broken down into subcomponents, sign detection, pedestrian detection/object in front of the car detection - and then there's logic that encapsulates these classifiers using them as inputs to determine how best to steer and power the cars drive wheels.
Its a bit far fetched to believe programming has fundamentally changed, at least not yet
He also points that the challenges are not where academia is focusing.
The car is parked IF it is on the side of the road AND it hasn't moved in X time, AND ... But not if... etc.
Software 2.0: The car is parked if the neural net says so.
Okay, if software 2.0 is all about thinking at a higher level, and training the neural net to deal with the details, why is the focus on detail like "is the car parked", or "is it raining", or "where is the lane marker?" Why can't we train "this is good driving" / "this is bad driving".As a software 1.0 programmer I can see how that seems completely unreasonable, but it does seem to follow the logical direction of the talk.
Most people also do some experimentation and calibration. Check how much empty space there is after parking, drive a circle on a snowy parking lot until you spin etc.
He did talk about the problems of complex models. They mostly treat the models as fairly fixed (see piechart slide of PhD vs Tesla). Most of the challenges are in labelling data.
See here [2] for an example of production ML testing practices. I wonder how much of this is in place at Tesla? I would argue they should be at the forefront of work like this. Something tells me they aren't.
Prodigy is an annotation tool that makes it easy to use active learning or other model-in-the-loop features. It's a downloadable library, that can start the web server on your local network, allowing 100% data privacy. We've just rolled out experimental image support in v1.5.0.
Fundamentally, the problem with these systems (and note, sometimes we say this about people too) seems to be a failure to think logically. Perhaps expert systems with sufficiently detailed logical data sets could enable more complex frameworks for decision making, and allow systems to dynamically create and run judgment calls with the NN classifier IDs and confidence levels as input sources.
The only bits added on are references to Tesla.
If all that we get out of it is fancy data labeling tools, incapable of learning anything new by themselves, it’s going to get old real quick.
For example:
- automatically learning that trolley is a not a great fit for a "car" because they are behaviorally different
- reclassify that cluster as a new entity even if it doesn't know the english term "trolley".
- If it finds the new distinction useful, ping the human that it needs more training examples for that situation
Similar to how a human learner can identify that he's bad at something, figure out what the common problem is, and use that information to focus on what to practice on next.
I am sure he alluded to doing this in his talk but what's the technical term for it?
For example in the bright smudges vs raindrops case, it might not have a label for the sun but it should be able to identify it as a important dimension in the cases it is getting wrong. Better yet, something more abstract like "smudge illuminated by light" or "bright background" that will be hard to annotate (e.g., how bright is bright?).
There are some practical software issues around not knowing the number of classes in advance, but those are "just coding".
There is no reason why introducing a new class shouldn't be as simple as providing additional examples to an existing class.
I thought their approach to rain sensors was interesting. The vision AI wiper function seems like overkill when a different system is capable of performing it almost flawlessly. I'd guess that the AI has a dedicated circuit though and becomes upgrade-able.. so those are pluses. But the rain AI system is a good test case and learning task for both the humans and the AI. Hypothetically, if you can't recognize rain drops, how can you recognize cars? It sounds like they learnt a lot trying to make that function so hopefully a lot of the knowledge in building that system generalized/translates over to the rest.
Besides that, modularity is an important design principle. It would be interesting to see how people combine different NN modules and integrate them with 1.0 code. Do you have a NN 2.0 controller? Some kind of self learning system that you train? I would imagine you'd want to take feedback into account at some point probably in Bayesian way.
Also an interesting take on how complexity has shifted from architecture selection to labeling data. It makes me wonder if there won't be a "Software 3.0" where most of the complexity shifts from creating a good labeling schema to, say, deciding on a good evaluation metric (I think the buck stops here as I can't imagine an AI silver-bullet automatically determining the evaluation metric). Perhaps unsupervised learning will come to the rescue and free us from the complexities of label schema design.
Have two or more labeling teams labeling the same stuff so you can reach a consensus or flag the differences and review and figure out why there was a difference.
Humans will be doing this for a while, I think it's worth having large companies (as large as Goog/Amzn/MS/FB) dedicated to the task.
Given the shown technique to build a state of the art neural net, I wonder what QA will look like and if we will be able to reach a sufficiently low probability of failure.
5 sigma reliability will be necessary at least in some fields for humans to accept to rely on it (like autonomous driving)
What the hell? 52? Really?
Software 2.0 is translating human intuition into machine code directly through advances in machine learning. How well someone’s dataset has been labeled will determine how well a “software 2.0” program will work, since that’s where where the human intuition lies.
Fortunately, I now have a fairly refined method for managing data, labels etc...
Google tried end-to-end deep learning AV systems and failed exactly because of the reasons he went through at the end of his talk.
You can't debug a neural net
Obviously not, but you also cannot hand-write code to do what neural nets do, so what's your point? If you can make your neural net 1000x better than a hand-written algorithm, or if Tesla Autopilot is 1000x better than human drivers, that doesn't matter. It's not "playing with human lives" if the humans around the car are 1000x times safer than sleeply, distracted, or violent human drivers.Set of all computer programs has two subsets.
S = { s | Static Analysis can be performed on s }
N = { n | n makes use of a neural network } with N ⊄ S
Let n ∈ N and s ∈ S. There are a certain set of programming tasks T = { t | t can be solved with n but not s }
Thus any claim that using n to solve t is "unsafe" because you cannot perform Static Analysis on it is absolute BS, because programs s ∈ S can't even solve the damn problem!For example, how do we know that a task solves a particular problem if we can't perform static analysis on it? It may give the appearance of working and then degrade radically under certain conditions. That really matters if you're using it for safety-critical applications and the problem space is large enough that it can't be exhaustively tested.
I mean, there are definitely huge shortcomings in the "let's have the computer build the model from examples" approaches, and some of them are talked in the video, others are not :
- rare events are hard to train (that's talked about). The problem is that it's a long tail of unusual events.
- models generated can't be statically analyzed. You can't predict what's going to work and what's not. You can only hope. One very striking recent example is in this video : https://youtu.be/w2BWmSBog_0?t=220 . Here you can see that an AI trained on the model of AlphaZero managed to reach 3223 elo rating (so, far beyond human), yet it blundered its queen. And that's just chess, where every rules are written in advance.
- Models don't build human knowledge. That's more of a philosophical point, but imagine a perfect AI built on neural networks after having read all human knowledge. What can it teaches us ? AlphaZero chess isn't able to provide any clue or explanation on why it favors one move instead of another. You can only learn buy playing against it, but that's all. Not even the developpers can tell you what advances in chess theory has AlphaZero made.
It's a road system designed for humans, and with current tech, you'll always have issues chasing the long tail down. Though I'm fairly sure self-driving is still statistically better than humans behind the wheel, especially distracted humans as the information age has made us.
The more immediate solution is white-listing safe roads that have sane paint lines and are relatively straight-forward roads. That way, the chance of the AI getting it wrong is drastically reduced. This covers most highways where you get the most out of self-driving systems anyway. Bonus points if you can automate this decision-making with recent satellite imagery.
The need for lines is only a temporary crux. AI systems will continually improve until the white lines are guidelines only. There are lots of other factors in an image that can be used to deduce where the lines should be.
But there will have to be handling of uncertainty which humans have to do too, simply just slowing down (not sharp braking) will help most cases.
Another analogy could me made to data entry: you could allow free-form text input and try to process all of the long tail of unexpected inputs, or else set out a format that the data has to follow and enforce validation at the input.
- Humans also do unexpected things, like stepping on the gas instead of the brake. If Autopilot does unexpected things at 0.01% the rate that humans do unexpected things, then it is a huge safety bonus to use autopilot.
- How can we solve image recognition any other way? We must move the needle forward on our technology. If we do not struggle against the adverse side-effects of our software and make it better, we will never advance it and we will be stuck requiring human drivers for all driving tasks.
You’re forgetting a variable - the frequency or likelihood of a situation to occur. Go far enough down the long tail and neural AI can get far more deadly than human drivers.
Your point was that mosern autonomous driving systems can drastically reduce fatalities over human drivers—-and I agree, but only for circumstances for which the car’s neural systems have been well-trained. But the systems ought to be able to handle ~99.99% of the types of circumstances gracefully before most of us will trust them to safely drive us around.
But the difference is, one human doing the wrong thing does not mean every human Will do the same thing, given the same scenario. But whereas a software running on all cars will exactly do the same wrong thing given the same input. Sure, three is also an advantage (arguably), as fixing it once fixes all cars but that has other problems in taking on the update.
With AI we could enter a world where you'll have 0.0001% of having an accident, but this could happen anywhere anytime in any random situation (such as a specially shaped cloud in the sky).
This is what makes it unacceptable, IMHO.
- Using data augmentation to turn the smaller amount of examples into enough samples for appropriate representation within the dataset.
- Add a weighting coefficient to the model's cost function to make misclassifying these examples more expensive.
Note: you can do serious harm to your model with either of these approaches if you don't know what you're doing. The safest solution is to collect more examples of the infrequent class.
I believe the last three directors of autopilot have quit in the last 3 years or so. And that's in addition to the mass exodus of executives, some of whom left millions of dollars of stock options on the table.
Talent is more pricy in places that lack them. Waymo can tap into googles vast amount of talent pool, this talented person would be worth less in waymo for sure.