TensorFlow Graphics: Computer Graphics Meets Deep Learning
medium.com
medium.com
There's also some approaches like SUGAN (https://machinelearning.apple.com/2017/07/07/GAN.html), but those don't really use differentiable rendering -- they just use traditional rendering, then apply a 2D GAN on top to make them look more realistic (from a CNN's point of view, at least).
This case is a little bit of the reverse, in that it's focused on making the computer vision component (the discriminator) try to match the visual content that has already been generated.
Seems like these types of "adversarial" approaches will be used in lots of different domains, as so far they've produced some pretty amazing results.
(If so, it should probably be done with multiple reference photos from different angles, to ensure shaders and lights aren’t adjusted badly so that the scene only looks good from the one angle that the computer was looking from when it was tweaking.)
Couldn't we just automatically generate lots of graphics renderings (that are by definition already perfectly labelled) and use them to train a ML model? That wouldn't require any differentiability, would it?
So how is this approach different?
It allows you to get from "these 1000 output pixels were wrong" to "the truck in the scene model actually was two inches to the left compared to what I expected/predicted" or "the texture of that apple should be changed this way to match reality".
You wouldn't use it to tweak labels for image classification tasks, but to learn better underlying models of physical reality and behavior.
Is the goal to tweak the rendering parameters until it matches a given input image? If yes that would be inference, not learning. What am I missing?
Edit : nevermind, the article does state that it is similar to an autoencoder
No. The sentence of the parent post "a pixelwise error can be computed and backpropagated to the CNN" is possible only if the renderer is differentiable.
Update:
It would require more training cycles and would not be as "atomic" as iterative tweaks but seems possible.
At same time, I wonder about making loss function talk with some external renderer would make it possible to mix both approaches.
I suppose you theoretically could do it with some trial and error method or grid search or something like that, but it's going to be absolutely computationally unfeasible in the general case; the pixelwise errors only become practically useful if you have an uninterrupted differentiable/'backpropagatable' path from your parameters to the pixels.
There are other use cases, for example, rendering "what-if" scenarios for reinforcement learning - e.g. you may have a model that predicts that if an agent does action sequence ABC then it will result in a world state X, and it can render an image Y which it expects to perceive. When it actually performs these actions, it actually obtains a different image Y' ... and needs to update its prediction, so it'd need differentiation/backpropagation through that rendering (and beyond) to figure out how it could have predicted the correct world state; so this feature is needed to allow learning the predictive model.
Seems like a nice way to debug and create more complex models.