Event-based camera chips are here, what’s next?
spectrum.ieee.org
spectrum.ieee.org
I think they're going to be used a lot in robotics in the future, it doesn't seem to make sense to take a bunch of frames, several million pixels each, and spend a lot of compute power to find features when consecutive frames are so similar.
You and I don’t see anything like pixels; our vision system pulls out relatively high level features like parallel lines, motion etc and starts working with that in addition to other optical data, muscle feedback on your lens shape etc. This is why many animals, including humans, freeze when frightened or when they perceive a risk: stationary objects are simply harder for most animals to pick out. The theory is (or was — I don’t know what they are doing today) is that you can see things interesting to humans and reason about them in ways interesting to humans.
Not part of that company: there are some wild but interesting theories related to this kind of thing. We pretty clearly have face-finding hardware (though not in the early visual system I presume). Do animals have this? Cats and dogs look at human faces.
I have read a theory that reading may hijack some subsystem originally used for identifying tracks (footprints).
It does appear that recognizing 2D pictures is learnt, not innate, while 1990s/2000 ML went the other way.
Some anthropomorphically-inspired design could yield systems that are more comprehensible and more useful. I wouldn't make a fetish of it (cars don't run faster than a human does) but machines should be better adapted to humans rather than the other way around (e.g. cars need special places to move around in so can only help humans when that is possible).
Yes it seems we do not see and think in pixel, but are we really not think in pixel even though we are not aware of them ?
If you look at the anatomy of the retina or the physiology of the early visual system you’ll see that the concept of “pixel” doesn’t even really make sense in that context.
We don't see the pixels because our consciousness works at a higher level of abstraction. You won't find pixel a few layers deep into a CNN either, but that doesn't mean they don't exist, the "neurons" at that layer just don't see them because they have no use for them.
That's always been my understanding at least...
>[Car makers want in-cabin monitoring of the driver to ensure they are attending to driving even when a car is autonomous mode.]
Automotive EE, I can tell you with certainty what our customers do not want. It seems pretty obvious to them that their images will not stay in the vehicle, and your car will snitch on you.
It's been available in cars for 15 years now, and is used by a lot of manufacturers.
> It seems pretty obvious to them that their images will not stay in the vehicle, and your car will snitch on you.
That would be against European law:
> Driver drowsiness and attention warning and advanced driver distraction warning systems shall be designed in such a way that those systems do not continuously record nor retain any data other than what is necessary in relation to the purposes for which they were collected or otherwise processed within the closed-loop system. Furthermore, those data shall not be accessible or made available to third parties at any time and shall be immediately deleted after processing. — regulation (EU) 2019/2144
Your company will probably offer this feature soon, if they don't already do.
https://en.wikipedia.org/wiki/Driver_monitoring_system
https://www.aptiv.com/en/insights/article/what-is-a-driver-m...
This seems like one of those "obvious in hindsight" discoveries, which are always the best ones.
It's only really in the context of computer vision and object tracking that the brute force whole-frames model starts to seem less than convenient.
Yet, this also means the possibility of some missed changes, of some changes being taken as lighting changes, instead of object change.
Was that a cloud, or a large shape close to the lens?
Hmm. Gonna have to read on this.
Normal cameras lose information by accumulating the intensity of each pixel over time until the next frame "arrives". You can not reconstruct that information anymore by diffing the frame-based output.
Event cameras on the other hand track the timestamp of the intensity change for each pixel individually. Thus they experience far less motion blur because they don't average the signal amplitude over time.
However, I haven't yet understood why their dynamic range is also better.
If you’re sampling intensity changes at high frequency, then you don’t need to worry about saturating each pixel between each sample interval. Instead you would need to integrate over all the collected deltas to get an intensity value for a given time period.