Reflected hidden faces in photographs revealed in pupil
kurzweilai.net
kurzweilai.net
Ah Moore's Law. It explains everything. Double every twelve months, eh? Hmm, quick math. Facebook has been around for 9 years... oh yea, that explains why Facebook allows you to upload full resolution 60MP photos. /s
>What do your Instagram and Facebook photos reveal?
Nothing. They're about 100 times too small for this technique. And they probably always will be.
Sure more pixels usually mean slightly more information, and there's obviously techniques: stability-enhancement, faster shots etc, that change this critical point. But for the average person carrying a camera: we crossed the critical point long ago and now just waste a lot of SSD space.
This results in an obvious merging of the blue/green strips. The more the diffraction the more they merge, the less sharp the images become.
Current APS-C sensors at 18 megapixels have pixels small enough that an aperture of around f/11 is the minimum size before this effect begins. As sensors get more and more dense, the required aperture size gets larger and larger. F/8 is considered a 'normal' aperture and once sensors hit that density then the payoff from increased resolution is significantly less and can impact image quality adversely.
This is why science missions use 2mp CCDs that they know the characteristics of, rather than some 40 megapixel phone sensor.
edit: I've just realised I've typed all this out with the wrong idea. You wanted a simple explanation of diffraction. Imagine a water tank with waves being generated from a source at one end. Put a wall with a narrow hole in the middle of the tank. If the wavelength of the waves is significantly less than the width of the hole, they will for the most part pass through unaltered. As the hole gets smaller, more significant changes to the waves occur. They begin to 'spread out' or 'diffract'. An intuitive explanation is difficult but this is a practical one that makes sense to people.
There's also technological barriers for creating large megapixel CCDs as opposed to CMOS. Large CCDs are usually created by "stitching" multiple sensors together, which increase their price as a square of the size.
You can shoot fullframe (35mm) camera at ISO 25600 and get decent results, but if you crank up signal amplification on iPhone to 25600 you'll get nothing but random color dots :-)
It would be foolish to suggest there was a single motivator for inclusion of a sensor. In reality it is a combination of our answers. Large sensors allows lower gain and better noise performance, enables smaller aperture use at acceptable sharpness, takes less resources to process, transmit etc.
You're not wrong in mentioning it though. Although FF at 25600 is going to be piss poor even if it's Nikon / Sony.
edit: corrected stupid misspelling
Nikon D600 at 25600: http://www.imaging-resource.com/PRODS/nikon-d600/nikon-d600G...
Not even close to "piss poor" in my book.
That being said, you don't go that high unless you're very desperate and the sensor itself can go up to 6400, the rest is software.
That was f/32. The rings around the street lamps are diffraction artifacts.
As a result, higher pixel densities will cause light from a single point to be spread across multiple pixels, with different wavelengths going to different pixels. Software can probably reconstruct the image somewhat, but I'm sure you'll hit a limit where there is too much interference to get the image any sharper no matter how many pixels you squeeze in.
What you've described is chromatic aberration.
Diffraction refers to the spreading of waves around an obstacle -- such as the aperture blades in your camera. No lens need be involved -- you can get diffraction with a pinhole camera.
The higher your resolution compared to your aperture, the less photons you get per pixel.
Sensor size doesn't matter at all.
Diffraction cannot be solved as easily.
If your sensor counts a mean of 100 photons per pixel, then you'll see shot noise with standard deviation of 10 photons in each of those, for a signal-to-noise ratio of 10. If you quadruple your pixel size and now measure 400 photons per pixel, then your SNR goes up to 20.
This is why bigger sensors (more captured photons for same light level and exposure time) are fundamentally better at image capture.
I realize that is a very flawed analogy but it's the quickest I could come up with.
Here's more: http://en.wikipedia.org/wiki/Shot_noise#Optics
He never said that Moore's Law explains everything. He simply said that Hendy's Law explains why phones will soon has 39-megapixel cameras. Since the Nokia Lumia 1020 is available today with a 41-megapixel sensor, I do not find this to be farfetched.
Facebook and Instagram will depend not on Moore's Law, but on availability of bandwidth and cost of storage. To argue against Moore's Law for Facebook is to argue against a strawman. The article did not make that connection -- you did.
In 2005, would you have predicted that Youtube would never carry 1080p videos? I mean, back then, Blu-ray and HD-DVD were still battling it out for dominance of the next-generation disc market.
I was referring to the fact this is on kurzweilai.net and Kurzweil tends to think it does explain everything: how we're going to beat the turing test in 16 years, how he's going to resurrect his dead father, etc.
From the subtitle to the last paragraph, the article clearly implies that this technique will soon be applied to social networking sites. And it seems to use Moore's law to handwave the details.
Famous last words when predicting non-advance of technology.
Please help me understand a future where social networking sites will maintain and display ~ 6200px X 6200px photos of your selfies. Outside of genetically engineering new eyeballs, I'm struggling here.
I think it's safe to assume that FB has tried to keep the original version of every photo ever uploaded. A few years ago, YouTube retroactively upgraded uploaded videos to HD. It's also safe to assume that they're going to continue upgrading/evolving/extending their most important feature, photos, and try to make it the #1 place to organize, store, and share all your photos.
Facebook then becomes an API upgrade away from exposing the original versions of uploaded photos. The photos will also come with all of your friend's faces conveniently tagged, allowing an app with the right permissions to extract high resolution faces/pupils/reflections. None of this needs to happen in the news stream.
So they may not have full resolution versions from the mobile, which is likely their primary source of pictures.
For a while, the Java upload tool was knocking down resolution before upload on desktop, too. That was a few years ago, I think. 2008-09?
I'm not saying you're wrong. But look at it this way: you are predicting that the resolution of stored photos won't increase by another power of ten. If you are right, this would be the first time in tech history that such a prediction would be correct.
Also, just a nitpick. I think relevant factor is ~100, the difference between the current resolution maintained by most websites and the proposed one for this technique to start becoming feasible.
[Ok, 20 may be right. I didn't know they allowed 2000x1000]
Personally, I will enjoy seeing high res pictures of other people on my 4k monitor when I get one.
Again, not saying I'm 100% that you're wrong, but you're making a bold and unprecedented prediction.
it would be cool to make group pictures, though.
10 years ago, I never would've guessed we'd be up to 120" TVs with 4k resolution. I was pretty happy with my 42", which was a big step up from what my parents had used before.
Right now, those examples I mentioned are new and expensive, but it won't be another 10 years before they come down in price and Sony and Samsung will be looking to sell us on another resolution.
Meanwhile, having computers hooked up to your TV is also becoming increasingly popular. My dad hates looking at a tiny computer screen, and would much rather look at photos on a TV.
So you've got high resolution TVs, and low resolution photos. That means the photos are too small or blown up to show their pixelation. That won't fly forever.
So, your limiting factor is glass, and manufacturing precision on that end is moving forward much slower than on the sensor.
[1] - http://www.luminous-landscape.com/tutorials/resolution.shtml
I wager we'll see 4k tablets under $1k within 2 years, and TVs will have to follow suit (drop in price, increase in quality).
Now imagine that lenses for single-shot panoramas are standard on all cameras and cellphones.
I used to think my 50GB disk was massive. Now I have more RAM than that, and use it all!
Photos will increase in size because of many factors, possibly including multiple focus points, or more color information, or 3d or a whole suite of yet to be thought up reasons.
I am very sure social networking sites will store larger files than 6.2k x 6.2k, with the 4K monitors and video and cameras producing larger and larger images (think medium format, even a piddly little smartphone now has 49MP!)
This saves a big resolution picture with spacial info. With your computer then you see a smaller part of the image. 6200x6200 images is not too much.
Add everything that has to do with 3D. Life is 3D, not 2D.
I can see the appeal of life-sized photos on social networking sites in the future. Group photos could benefit from being larger than that to include multiple people at life-size, but I'm unsure of how a resolution greater than 2400x3000 would be useful for a profile picture at this point.
[1] http://upload.wikimedia.org/wikipedia/commons/6/61/HeadAnthr...
[2] http://prometheus.med.utah.edu/~bwjones/2010/06/apple-retina...
A few people are pointing out use cases where a 39MP photo might make sense down the line:
-Group photos
-Head to toe shots
-Panoramas
Let me just point out that this research is done on "high-resolution passport-style photographs." That is a close up head shot with a normal FOV. You fit your whole body, your friends, and everything else in the photo... you're going to need hundreds to thousands of mega pixels for this to work. At the end of the day, it looks like you still need > 39MP just for the face.
"The room was flash illuminated by two Bowens DX1000 lamps with dish reflectors, positioned side by side approximately 80 cm behind the camera, and directed upwards to exclude catch light"
Yes, that's the kind of lighting conditions that are routinely present at crime scenes.
Not that this isn't interesting, but their hypothetical (hostage situation) would have been a much more intriguing test case. But, since those pictures/videos are rarely taken with super high resolution cameras with good lighting, it's unlikely this technique would be helpful.
Edit: Just to be clear, I think anything new in the form of research and development is generally good. Maybe this work will be built upon to create something great. I just don't like the implied promise that pieces like this make. It oversells what technology can actually do and is, I think, misleading.
So CSI has no excuses still.
Imagine what happens when we have several orders of magnitude more computing power available. A lot of things aren't possible today because they are too computationally intensive, but computing power is just getting cheaper and more abundant. Imagine what happens when it's feasible to process basically all data on the entire internet. Every photo and video. Every tweet, vine, and wall post. Every porn video and macro image.
It'll be possible to find the identity of anyone in any picture. Moreover, every picture or video that someone has been in will be easy to correlate. Now scale that up to every activity. Anyone with enough computing power will basically be able to create a dossier on anyone, or on everyone.
The implications are concerning.
I'd assume that most people have been facetagged in enough photos on Facebook so that they can be found in every photo by every random person who uploaded there, or to a few other cloud services, which any reasonably determined organization can scrape. People photograph a lot.
And then the CC cameras as well.
Though there are easier ways to spy on people, you can get all relevant info from their phone apps and social media services anyway.
For example, consider identifying people by the way they walk, the clothes they tend to wear, their body build, etc. That sort of thing makes witness protection efforts practically useless, for example. The sheer amount and detail of information that could become available will enable things that we can't even remotely imagine today, but ultimately the biggest change will be loss of anonymity at every level and loss of privacy.
If you can analyze video at, say, 10,000x real-time you can crunch a billion hours of video in a decade. Imagine the value of the meta-data you could get from such an analysis. You would have incredibly fine-grained details on people's lives, especially as video use becomes more common. Being able to see hidden details in reflections is immaterial compared to what will be possible. With enough computational resources it'll be possible to create dossiers on everyone everywhere. You'd know where they live, who their friends are, what they do for fun, what they own, what they wear, what they read. Using writing analysis you could find out how politically active they are, what their opinions are, and so on. Given that abundant individual liberty is still a fairly rare, and constantly endangered, condition on Earth it's frightening to imagine tools that could give such information in the hands of unscrupulous folks.
100,000,000 / (24 * 365) = 11,415
313 million people in the US. Computing clusters of similar capability required to observe them all:
313,000,000 / (100,000,000 / (24 * 365)) = 27,418
Granted, better or worse depending on how much you want to look at of their lives. If people only film important events is that better or worse for you? If they film important events how does that shift the balance of evidence from what the video contains as compared to the mere fact that they filmed the event?
I think... I fear that in many ways what you're talking about - video processing aside - is already here. If I put my evil hat on for a minute:
Okay, we want to work out what someone buys, where she lives, what she wears, who she knows, where she goes...
Where's that information already written down?
• If she went into a store her phone probably pinged the wifi, whether or not she connected:
http://www.tomsguide.com/us/Wi-Fi-tracking-Nordstrom-Cisco-R...
• ANPR systems are already widely deployed, so that's anywhere she went in a car:
http://www.projectcensored.org/police-tracking-license-plate...
• Credit card companies keep records of your purchases, obviously.
• There, probably, are going to be pictures of her on facebook.
• Facebook can be easily mined for relationship details - most people don't set their pages to hide their friends lists, despite the fact they should.
• Your IP address and browser fingerprint is trivially trackable.
For comparison you need 32.7 bits to pick someone out of the world population. Doesn't make you uniquely trackable of itself but puts you in a fairly small group.
... And the topic and language someone chooses to write about online... well, I agree but I don't think you need particularly powerful computers to do that - and given that I suspect there's more powerful evidence out there I'd expect that only to be deployed against individuals of interest.
#
I don't think the big threat to privacy is more powerful computers. Though they certainly don't help, I think the damage in that sense was done with the first ANPR system. I think the big threat is networked databases that already exist and take much less effort to extract relatively strong evidence out of.
I've also got some reservations as to what you're talking about is comparable in complexity to just encoding the video for upload. I agree that that is processing, but the data's already structured in a known way and the transformations performed on it are relatively simple. The question to my mind here really does make Moores law the important thing. I know that youtube does try to check for copyrighted content - but equally they're not particularly good at it even there.
Unfortunately I can't really visualise exactly what you're talking about in my head in terms of what you want to do to the data - face-recognition would presumably be part of it, as would voice to text, and then you'd want to check that against databases of stored signatures. Possibly you could make the task vastly less complex by using other data around it to limit your database queries, but still. That's not a simple thing. Against individuals of interest, I could totally see it being done, but everyone seems very wasteful - I'm not sure how the small amounts of evidence I think you'd get would justify the costs unless costs became very low.
If we're looking at something like a 20 magnitude increase in computing power over the next fifty years or so, (and similar increases in access speed and so on,) then all those concerns just sort of vanish away into the sea of massive power. ^_^
I'm not well read on the topic so I might be talking out of my ear, but this is the sort of thing I'm anticipating getting more and more powerful over time: http://users.soe.ucsc.edu/~milanfar/talks/milanfar.pdf (jump to page 23 if you want to see demos of the algorithm)
http://www.scriptol.com/programming/graphic-algorithms.php
edit: grammar
Especially in certain domains such as written letters (e.g. license plates) or faces that have a well-known structure with variations in known dimensions, I wouldn't be surprised to see algorithms that can probabilistically infer what information would have originally been there. Of course you'd never reconstruct an entire face from a single pixel, but how far can we go? I dunno - but my hunch is pretty far
wouldn't or would ? don't or do ?
Medium format cameras have a much greater area sensor and probably overall better quality ($2,000 vs $20,000 or even $50,000 medium format camera) compared to a regular digital DSLR camera. Even a quick Google seemed to indicate photographers prefer the larger sensor size even if it has fewer pixels.
The article mentions "face images retrieved from eye reflections need not be of high quality in order to be identifiable" but even so there must be a massive difference in picture quality compared to a Hasselblad medium format camera sensor.
Still interesting though it's as if the goofy CSI or Blade Runner "Enhance!" has come true.
Yeah, it would be interesting. But would you hold your breath for it? In an age where kids "shoot themselves" while being handcuffed and whatnot, I for one wouldn't.
This is not how one normally does science.
EDIT: Oh and camera-charecterising noise signatures.
The video [2] gives a good explanation, and the revelation of the playing card is what strikes me as similar to revealing reflected hidden faces.
And, I'm now waiting for someone to release a "Privacy Eye" Photoshop filter (a la the ubiquitous "red eye" filter).
Is this a new idea, or have people been working on this for a while?
http://www.flickr.com/photos/oatmeal2000/8545400550/sizes/o/
You can clearly see that he uses one umbrella to the right of the camera. With higher resolution images and better focus you can sometimes determine the brand of lighting they use, but that's only if it has a very distinctive shape.
http://www.cs.columbia.edu/CAVE/projects/world_eye/
"The World in an Eye," K. Nishino and S.K. Nayar, IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Vol.I, pp.444-451, Jun, 2004.