DifuzCam: Replacing Camera Lens with a Mask and a Diffusion Model
arxiv.org
arxiv.org
https://waller-lab.github.io/DiffuserCam/ https://waller-lab.github.io/DiffuserCam/tutorial.html includes instructions and code to build your own https://ieeexplore.ieee.org/abstract/document/8747341 https://ieeexplore.ieee.org/document/7492880
And yes, Laura Waller at UC Berkeley has been one of the pioneers in this research for a few decades.
The downside of this is that is heavily relies on the model to construct the image. Much like those colorisation models applied to old monochrome photos, the results will probably always look a little off based on the training data. I could imagine taking a photo of some weird art installation and the camera getting confused.
You can see examples of this when the model invented fabric texture on the fabric examples and converted solar panels to walls.
It inevitably means that the "camera" visually parses the scene and then synthesizes its picture. The intermediate step is a great moment to semantically edit the scene. Recolor and retexture things. Remove some elements of the scene, or even add some. Use different rendering styles.
Imagine pointing such a "camera" at person standing next to a wall, and getting a picture of the person with their skin lacking any wrinkles, clothes looking more lustrous as if it were silk, not cotton, and the graffiti removed from the wall behind.
Or making a "photo" that turns a group of kids into a group of cartoon superheroes, while retaining their recognizable faces and postures.
(ICBM course, photo evidence made with digital cameras should long have been inadmissible in courts, but this would hasten the transition.)
Sworn testimony is admissible in courts. I think the "you can just make evidence up" threshold was passed a few thousand years ago. The courts still, mostly, work.
But I'm less worried about the courts and more about media that might publish photos without realizing they are AI generated - or ordinary people using those cameras without understanding how they work and then not realizing there may be some details in the pictures that are plain fantasy.
interesting.
even if this does not immediately replace traditional cameras and lenses... I am wondering if this can add a complementary set of capabilities to a traditional camera say next to a phone's camera bump/island/cluster...so that we can drive some enhanced use cases
maybe store the wider context in raw format alongside the EXIF data ...so that future photo manipulation models can use that ambient data to do more realistic edits / in painting / out painting etc?
I am thinking this will benefit 3D photography and video graphics a lot if you can capture more of the ambient data, not strictly channeled through the lenses
Intuitively, imagine moving your eye at every point along some square inch. Each position of the eye is a different image. Now all those images overlap on the sensor.
If you look at the images in the paper, everything except their most macro geometry and colour pallet is clearly generated -- since it changes depending on the prompt.
So at a guess, the lensless sensor gets this massive overlap of all possible photos at that location and so is able, at least, to capture minimal macro geometry and colour. This isn't going to be a useful amount of information for almost any application.
The sensor has a zillion pixels but each one only measures one color. for example, the pixel at index (145, 2832) might only measure green, while its neighbor at (145, 2833) only measures red. So we use models to fill in the blanks. We didn’t measure redness at (145, 2832) so we guess based on the redness nearby.
This kind of guessing is exactly what modern CV is so good at. So the line of what is a camera and what isn’t is a bit blurry to begin with.
The wikipedia article on demosaicing has an algorithms section with a nice part on tradeoffs, how making assumptions about the kinds of pictures that will be taken can increase accuracy in distribution but introduce artifacts out of distribution.
The types of models you see used on camera are pretty constrained (camera batteries are already a prime complaint), but there’s a whole zoo of stuff used today in off-camera processing. And they’re slowly making they’re way on-camera as dedicated “AI processors” (I assume tiny TPU-like chips) are already making their way into cameras.
Would love to lose the camera bump on the back of my phone.
Or no photos anymore?