I've noticed that your project for fast-neural-style does instance normalization over batch normalization.
Batch normalization has the benefit that you can merge the gamma & beta into a convolutional layer on the forward pass, which makes it a lot faster by allowing you to skip a step when building the styled images using a trained model.
Can the same be done with instance normalization? I didn't see a formula in the paper but I would think so, since they are fairly closely related.