(DSLR cameras have independent phase detection arrays. This is why the mirror has a small mirror behind. To illuminate that section).
Lenses are tuned. They know how distant the thing they are focusing on, also photos' EXIF generally carries the focus point information. Side note: Apple Aperture used to be able to show where you focused with that info. Still no app does this. I'm still mad that we don't have Aperture anymore. Anyway...
So, if you collect this phasing information alongside the photo which you're taking, you are capturing the depth map of the photo. Its resolution will be lower, but not lower enough to be useless.
For example, Sony (and most probably all other big camera manufacturers) cameras use this phasing data throughout sensor in real-time for following moving subjects to keep the focus on them by predicting where they are going.
So, there's no "in theory". The phasing data is the depth map. Otherwise your camera can't focus on anything. It's there to focus, and it's done by reading the phasing data and knowing where to go by how much.
All fast AF systems are PDAF. When light is too low, then it's CDAF, which does way slower, without using any phase/depth data.
[0]: https://en.wikipedia.org/wiki/Autofocus#Phase_detection
Edit: Meaning clarified, bugs squashed, mirrors polished.
A depth map from the apple camera (again, signed) would show that the entire image had the same distance from the camera.
This is a pretty standard application of trusted computing and can be done entirely on the iPhone. A server would only possibly be needed for anonymization (while retaining key revocation capabilities if a key does end up leaking), but there are serverless ways to do even that (TPMs have supported these for a while now).
I think a big part of validation for things like these are just "could it have been modified since Z event happened", because Z was not something people paid attention to before.
Nothing prevents anyone from opportunistically pre-generating and timestamping millions of permutations of fake kompromat and then selectively revealing the one that turns out to be useful after the fact.
You could charge per attestation, but the economics of that don't look great; you could demand publication of the image itself before attestation, but that would obviously not fly for most use cases out of privacy concerns.
If you're doing server-side timestamping and someone is sending millions of items, I think you just ban them. Apple accounts are free but not inexpensive.
There are already pretty good depth map generation algorithms (for a decade or so) which works on 2D images.
Generate the image, feed it to a depth map generator, viola.
However, you can't get it signed by the sensor itself. That'd be hard.
It helps but doesn't solve credibility issue.
Then we might move to an age where photos which were not signed by someone will be treated as fake.
What we need is almost something like the classic "press C+A+D to log in" or the idea of "above the browser pane"... something an image cannot show you unless it's legit. Perhaps a personal thumbprint icon that is different for each viewer, so you know its your device telling you that something is real and not just some mock fake thing?
It may not be in reach for you and I - but within reach of anyone with the budget to run a 'bot farm'
Proof of captured image with hardware-based-attestation can be increased by adding extra information besides the RGB channels, like depth mapping (which a projected screen wouldn't be able to fake), and eventually light field recording (plenoptic imaging, e.g. Lytro).
I suspect it’s significantly less than millions.
things like color rendering, sharpness, screen door/morie, motion blur, and yup even good ole depth queues in the lens system all show up as artifacts.
case in point, outside of a few specific niche cases like the mandalorian, that virtual production video wall thing is not actually being used all that much because of the amount of post shoot cleanup required to fix all those issues, it wasnt actually that much cheaper and its not really better either. especially when you factor in how hard it is to shoot that way.
Pre-AI fake videos sidestep this sort of issue by lowering the video quality.
Add some motion blur, some camera shake, some poor lighting, the camera being slightly out of focus, and make the video 720p instead of 4k.