It's seductive to believe that if we just use a clever model, we can create something that looks real.
It's not that simple. By definition, the closer that your engine looks to real life, the less flexibility artists have. If a scene looks perfectly real, artists would not be allowed to change anything at all, because any change would make the scene look less real.
Therefore, no matter what kind of clever mathematics you use, or any algorithm you come up with, your flexible art pipeline will torpedo your ambitions of creating a CG video indistinguishable from real life.
This is a fundamental limitation that I don't think has been fully appreciated, or at least isn't emphasized in literature. I think most people can't really believe or accept that no one, anywhere, has ever successfully created a CG video of a complex scene that is capable of fooling all human observers 100% of the time. (If a set of human observers are asked "Is this video real, or computer-generated?" then their responses should be no better than random chance.)
Your research looks promising! It would be interesting to pair an analytic solution with some hypothetical art assets that were somehow generated from reality, or otherwise fully capture all of the variables in a real-life object (i.e. the textures are more than simply RGB values).