How Self-Driving Cars Work
nytimes.com
nytimes.com
This article also misses a lot air companies like Daimler who showed off self driving car captibilities publicly in the beginning 2000 and on the web 3 years ago. It misses the prototypes shown off by Audi and BMW. It misses also the prototypes shown by Japanese car manufacturers.
From technology companies it misses Bosch who will show off at CES their own developed car and as such their expertise. It misses Delphi who works with MobilEye on self-driving platform to be used by the OEM, and it misses other TIER1 like Conti who have also shown their work.
By the way, the Mitsubishi Outlander is currently the best equipped series car who can do self-driving capabilities, it has a stereo camera system and Lidar combined with other other radar sensors. Next are the luxury cars from Audi, BMW, Daimler, and Volvo. Tesla here is last with only a simple radar and camera system (no stereo view), which has also no night vision capabilities like the others have.
They definitely have more than "just a processor": http://www.nvidia.com/object/drive-automotive-technology.htm...
I have a friend who works in the self driving team at Nvidia in Santa Clara (apologies for the login wall): https://www.linkedin.com/pulse/full-stack-engineers-self-dri...
They've also got some demos up on YouTube: http://youtu.be/YuyT2SDcYrU
So can I conclude there is a difference in the (state of) implementation between the hardware (sensors) and the software (self-driving intelligence) ?
Humans only have one sensor: a rotatable stereo camera. So at least in theory the number of sensors seems not the most important element ;-)
Tesla seems number 1 in pushing the frontiers in marketing, so it may also ahead in software, to compensate what it is lacking on the hardware side ?
I hear this comment a lot when defending Tesla's choices, and it's a red herring. The fact that humans only rely on two cameras means nothing. Repeating old comments of mine:
You also don't "need" megawatts of power to play top-level Go: humans do it with 100 watts of energy. Yet Google needed who knows how many megwatts of energy to train and run AlphaGo on their massive server farm.
Imagine two companies competing to win at Go, and one company had the attitude that megawatts of energy was not necessary for training and prediction, and another company threw the biggest GPU farm they could. The second company just played top-level Go this year. The first company is ~10 years away from a low-energy elite Go computer.
Humans implicitly perform SLAM (simulataneous localization and mapping). What do I mean? Look around your room. Close your eyes. Visualize the room. As a human, you've built a rough 3D model of the room. And if you keep your eyes open and walk through the room, that map is pretty fine-grained/detailed too and humans can keep track of where they are in the map.
Doing this accurately in moving environments (especially with lots of pure forward motion) with just two cameras is still a wide open research problem.
Doing this with LIDAR/GPS/IMU/recorded maps is solved. That's why people use LIDAR.
Matching the abilities of human perception is an insanely hard problem. Don't let cute problems like image classification fool you. Why make the problem even harder?
Visual slam is still linear-algebra/geometric/keyframe based traditional computer vision (including variants that incorporate GPS/accelerometer info). I think the state of the art is stereo LSD-SLAM, but I could be wrong.
It's also likely that AlphaGo is reasonably efficient, given that it used custom ASICs.
Low power is also still important due to issues of heat and energy availability. Low power also implies high efficiency which is important for several reasons.
The human brain is estimated at 20 watts (when people talk about computing systems they tend to not include the power needed for all the auxiliary infrastructure needed to keep it networked and cooled); it's also estimated that beyond 4 hours a day, learning effectiveness drops precipitously.
If we take the case of Go, you can take a 4 year old human and have a professional player by 13. This is about 950 megajoules spent by the brain while learning Go. For the machine, if you look at the learning part (self play, value and policy on 50 GPUs for several weeks) the estimate on energy spend is about 30,000 megajoules. The policy network is itself ~20,000 MJ, while the full AlphaGo system playing on a single GPU and 48 CPUs is just a strong amateur.
But this is not even an apples to apples comparison since the brain is not spending all of its energy on learning Go. In fact, learning how to play Go is very far from the most difficult thing the brain is learning how to do.
The computer can train itself with just records of past games and self-play. The world champion level human cannot. You must account for the difference in training.
To say AlphaGo or any RL system is learning from self-play is not in the typical understanding of the phrase. It's more akin to evolving with competitions against previous versions of itself, which should count as different instances. As stated on page 38 of 1604.00289.pdf
Between the publication of Silver et al. (2016) and before facing world champion Lee Sedol, AlphaGo was iteratively retrained several times in this way; the basic system always learned from 30 million games, but it played against successively stronger versions of itself, effectively learning from 100 million or more games altogether (Silver, 2016).
In comparison, from that same paper, it was estimated that Sedol could not have played much more than 50,000 games. My own estimate is about 40,000 games.
As for work required to learn, it's irrelevant to point out that one can learn from others. Learning, whether from play, books or study still requires energy spend and work by the learner. Most of the extra work is from study, occasional review with a tutor and discussions with peers--the last more of a meta-step: learning to learn. Accounting for books and some time with tutors will not, I argue, shift the budget much. Especially if you include that any machine playing Go requires overhead of power infrastructure, energy, cooling, networking equipment and occasional maintenance staff. And learning, improvement in architecture, requires searching through and discarding many changes and playing through a cumulative hundreds of millions of games.
The fact that humans have other available highly efficient means of learning is a boon and not a downfall. That's the whole point of getting to AGI. Learning from books and others is akin to learning by Program Synthesis from specifications.
The reason I noted the requirement of other trained professionals for training a human is that those other humans can distill what they have learned over years of play into simple rules. The machine can also use those rules, but the particular machine you are comparing to was specifically trained without any such rules and was required to synthesize them from scratch from historical games and self-play.
https://seekerblog.com/2006/01/31/the-murray-gell-mann-amnes...
For instance, a poli-sci major writing about a scientific/tech topic will likely get a lot of things wrong. However, the same journalist might likely have a lot of interesting things to write about international politics, because that's what his background is about.
that being said, I don't know how much journalist writing about politics have studied it in the past.
https://books.google.com/books?id=KUyOAgAAQBAJ&pg=PA42&lpg=P...
not that it really matters, i have an econ degree but wouldn't call myself an economist by any stretch.
the real issue, at least in the US, most news is really just either direct or indirect PR. curiously, i think 'X writer' is the term you are looking for. i.e., someone with a background in technology or fashion that reports on things is more likely to be called themselves a 'tech writer' or 'fashion writer'. it's almost like real experts don't even want to be called journalists. i certainly wouldn't.
i've found that you're much more likely to find accurate reporting in industry-specific media rather than general media, because then at least the sources of funding and agenda are pretty clear.
eg. Bloomberg, The Economist, The Wallstreet Journal, The Financial Times
Probably because it's the last industry where people are willing to fork over money for reliable news
[1] http://www.continental-automotive.com/www/automotive_de_en/t...
As a side note, how do you guys find out things like who buys who and whos's got what interesting tech? Especially when they're selling only to suppliers, as is the case here? You just stumble on the information, know someone who knows someone or what?
There are lots of companies working on low-end LIDAR. ASC's technology is known to work well; it just costs too much because it involves custom sensor ICs built with a nonstandard GaInAs process. Quanergy has been issuing a lot of press releases and getting funding, but they seem to have backed off from their < $100 solid-state sensor and are now selling $7500 rotating machinery like Velodyne's spinning top. There are also some companies touting continuous-wave systems, which usually don't work in sunlight or have much range, but are useful for indoor robots.
Fraunhofer is working on a technology for making flash LIDAR sensors using a regular CMOS process.[1] If that works, these things will come down to digital camera prices in volume.
Bloomberg has most acquisition information. I picked this up because I've been following ASC since the 2005 DARPA Grand Challenge days, when I went to Santa Barbara to see their technology. It was just on an optical bench then, not ready to deploy, but they had the technology with long-term promise. I dragged a VC down there to talk to them, but he didn't see near-term volume application. He was right; only now is there a market in sight for LIDAR units by the millions. High production volume is probably still 5-10 years away.
[1] https://www.ims.fraunhofer.de/content/dam/ims/de/documents/D...
[1] http://www.int-arch-photogramm-remote-sens-spatial-inf-sci.n...
The CEO of Mobileeye (who is an ex-machine learning professor) gave a very good, mildly technical talk on their approach at CVPR: https://www.youtube.com/watch?v=n8T7A3wqH3Q
The CEO of Mobileye is Ziv Aviram and he was never a professor. Amnon Shashua is the CTO, and he is still a machine learning professor at the Hebrew University.