Yes, it's a simple test:
Show a shuffled playlist of 20 videos, half of which are simulations. The simulations must be entirely synthesized; no mixing in footage of real life.
As long as the videos are non-trivial (~60sec and reasonably complex scenes), the viewers will correctly select all 10 of the simulated videos. If we were capable of generating simulated video identical to a camcorder, people would score no better than random guessing.
So not only have we failed to reach that goal, we are nowhere close -- it's effortless to identify all the fake and real videos. Especially nature videos, for example. When put side-by-side with real-life footage, our best simulation attempts are nowhere near the same ballpark. They're not even the same area code.
The idea that we really haven't accomplished this may seem offensive. My frustration is what led me to work on it. But unfortunately, we're nowhere close. I'm not sure traditional techniques will be enough to close the realism gap.
I've outlined a scientific test that anyone can conduct. The results may be discouraging, but they establish our capabilities. The goalpost is rigid.
To answer your other questions: Obviously everything is impossible until it's not. But I never meant to claim this was impossible, just impossible with our current techniques. I think physically-based rendering circa 2010-2020 will turn out to be a failure from the standpoint of replicating camcorder-grade video, but that's speculation. When you try to simulate physics in such great detail, you end up with hundreds of approximations. When these integrate together, you get modern Unreal Engine videos: they don't look real. They look pretty great! But not real.
Who cares about looking real? Well, like AGI, it's a goal worth striving for. Ever since we've been scribbling on cave walls we've tried to capture realism. Unlike AGI, there isn't much use for perfect realism; it might not affect the world at all. But I'm pointing at the summit; wouldn't it be a fun climb?