Sony Builds AI into a CMOS Image Sensor
spectrum.ieee.org
spectrum.ieee.org
I've worked with similar systems that used local AI to process speech (to determine things like turn taking in conversations) and there was a claim that the system enhanced privacy because no speech was ever recorded, but in truth it would have been easy to compromise that privacy protection. If the ability to record or export video is not part of the sensor design, then it would be difficult for anyone to alter the chip to record video, but how do you verify that? The chip could have a secret "test" mode where it exports video and AI parameters for troubleshooting. You'd have to trust Sony, which might be reasonable in some circumstances, but not in others. The same is currently true for a variety of phones with "smart" features associated with the camera. It may seem like the phone will only unlock for you when it sees you, but what if it also secretly unlocks when presented with a specific QR code? How would you ever know?
Definitely not going to stop state actors, but might help with some of the noise from doctored images/video.
You wouldn’t need to publish the unprocessed photo, just keep it in escrow.
The new photo would still get signed, and look original.
And if you did this in the right conditions, it wouldn't be possible to tell it was a photo of a print.
There are lots of demos of AI separating human voice from background sounds.
That said there are some types of processing that can improve audio quality. It's possible to remove some types of noise and to remove echoes from walls. You often need to know something about the room to do that though so there is no general "make it better" algorithm (yet).
One simple thing that sometimes helps is to adjust the equalization (increase or decrease gain at different frequencies), although if you use a normal equalizer this can add phase noise that makes the audio sound "muddy". A special kind of equalizer using FIR filters can adjust gains at different frequencies without altering the phase and results in a cleaner sounding adjustment.
There are "sonic enhancers" also, such as:
http://www.bbesound.com/products/sonic-maximizers/default.as...
which can add clarity to sound, though they don't work in all cases.
If you are mainly interested in intelligibility of voice there are algorithms for that which may destroy the musicalness of the sound, but make the speech easier to understand. Some of these things were first invented to improve the sound on the original telephones.
Probably the best way to improve sound is to improve the environment, often by adding sound absorbers on reflecting surfaces such as walls which gets rid of echoes. Even cheap microphones these days are pretty accurate at recording sound.
Quote 2: "..enhanced privacy.."
Syntax error.
When chip hardware gets faster, or you read about such developments as cramming more and more functionality onto the sensor, does that mean the package could draw less and less power and do the same functions as previously for less energy?
Suppose this chip's function were (as in the article) to image a scene and decide whether to send the image over wireless to a central monitoring station.
As the chips get more capable, does that mean less power is used compared to before, to do the same function? Or are there overheads that dominate the power consumption for such under-utilization of a chip? Is there a rapid falloff in benefit of the "advanced-ness" of a chip for such applications where you really don't need such sophistication, and rather have energy-conserving, simple design? Will this really lead to months more lifetime for a remote sensor powered by battery?
(sorry if some of my nomenclature is imprecise, but hopefully you get the idea of the question)
If it's doing new things like running inference (say, recognizing things using a machine learning model) it might take more power.
But sometimes it could use machine learning to replace previous functions and do it better with less power.
one example I can think of - most cameras have some sort of "motion detect" function. If machine learning could recognize people or cars or whatever is "interesting" better it might give fewer false positives, or transfer data less often. It might even use less power.
On the other hand, if it could detect people, and then further detect WHO they were, it might take more power.
* An audio amplifier: Power consumption is ultimately dictated by the energy conversion to sound and the conversion efficiency of the speaker. The system will consume at least that much irrespective of the technology developments. In practice however, the efficiency of the amplifier also comes into play. The static power drawn by an efficient amplifier would be much less than the above basics for creating sound, and so going to higher technology nodes would make much less percentage difference.
* Computations in a digital circuit: Power consumption in these is sometimes dominated by communication (interconnect; bus) like between processor and memory. If the bus is internal to the chip, advancing to a new technology node would typically result in power reduction.
* If the power is dominated by computations alone, there would surely be power reduction, again assuming the functionality is held constant. Circuits have static leakage power too, which happens irrespective of the computations, however, it is still usually lesser than dynamic power by design.
Typically however, more and more functionality is frequently added like in this case of computer vision added to the sensors, which results in an overall increased power consumption to an extent that the customers will accept (e.g., battery life asked for).
Back to this specific case of computer vision on the "edge":
Having the same functionality achieved by having computer vision on the sensor itself vs. the cloud, can result in orders of magnitude savings in power. There is a lot of power associated with transferring the data back and forth. Even if computations are handled by a processor on the same board as the sensor, significant power is spent as compared to when the communication lines are avoided.
https://gbdev.gg8.se/wiki/articles/Mitsubishi_M64282FP#About...
It's just a very restricted convolution.
Following are some links to such similar technology developments for those interested. Of these, [4-5] cover a lot more.
PS: I was the technology lead for [1] below, which I believe is a pioneering work which kickstarted this.
[1] Qualcomm Wants Your Smartphone to Have Energy-Efficient Eyes https://www.technologyreview.com/2017/03/29/243161/qualcomm-...
[2] emza Visual Sense - IoT Visual Sensors https://www.emza-vs.com/
[3] Why the Future of Machine Learning is Tiny https://www.oreilly.com/library/view/tinyml/9781492052036/
[4] The tinyML Summit https://www.tinyml.org/summit/
[5] TinyML: Machine Learning with TensorFlow Lite on Arduino and Ultra-Low-Power Microcontrollers https://www.amazon.com/TinyML-Learning-TensorFlow-Ultra-Low-...
AI is the keyword to get clicks.