That's not dismissive -- no one has ever made any program that outputs a string of images indistinguishable from a real camcorder. It's just that hard.
I think whatever the next leap forward looks like, it will come from a nontraditional approach. Something strange, like powering your real-time lighting model by an actual camcorder -- set it up, point it at a real-world scene, then write a program that analyzes the way the light and color behaves in the camcorder's ground truth input. Then you'd somehow extrapolate that behavior across the rest of your scene.
That last step sounds a lot like "Just add magic," but we have deep learning pipelines now. You could train it against your camcorder's input feed. Neural nets tend to work well when you have a reliable model, and we have the perfect one. So more precisely, you'd train your neural net against the camera's input video stream: at each generation, the program would try to paint your scene using whatever it thinks is the best guess for how the colors should look. Then you move your camcorder around, capturing how the colors actually looked, giving the pipeline enough data to correct itself. Rinse and repeat a few thousand times.
The key to realism, and the central problem, is that colors affect colors around them. The way colors blend and fade across a wall has to be exactly right. There's no room for deviation from real life. Our visual systems have been tuned for a billion years to notice that.
There are all kinds of issues with this idea: the real-world scene would need to be identical to the virtual scene, at least to start. The program would need to know the camera's orientation in order to figure out how to backproject the real-life illumination data onto the virtual scene. But at the end of it, you should wind up with a scene whose colors behave identically to real life.
It seems like a promising approach because it gets rid of the whole idea of diffuse/ambient/specular maps, which don't correspond to reality anyway My favorite example: What does it mean to multiply a light's RGB color by a diffuse texture's RGB value? Nothing! It's a completely meaningless operation which happens to approximate reality quite well. There are huge advantages with that approach, like the flexibility of letting an artist create textures. But if the goal is precise, exact realism as defined by your camcorder, then we might be able to mimic nature directly.
(Those dynamic occluders looked incredibly cool, by the way!)