FastPhotoStyle from Nvidia
github.com
github.com
- Their approach is a composition of 2 steps, what they call "stylization" and "smoothing".
- Top left of 2nd page they claim: "Both of the steps have closed-form solutions"
- Equations 5 is the closed form solution for the "smoothing" step.
My question: Where's the closed-form solution for the stylization step that they're claiming?
Are they calling equation 3 a closed-form expression? In this case the title and the claim in the introduction are rather misleadinng, because computing 3 requires you to train autoencoders.
I would still argue the term closed-form is misleading here, because:
- Even during training at any given time you can read off a "closed-form expression" of the neural network of this type, so closed-form in this broad sense really doesn't mean much. Furthermore any result of any numerical computation ever are also closed-form solutions according to this, on the grounds that they result from a computation that completed in finite number of steps. So really whenever you ask a grad student to run some numerical simulation expect them to come back saying "Hey I found a closed-form expression!"
- The reason the above is absurd is that these trained NN's aren't really solutions to the optimization problem, but approximations. So this is really saying I have a problem, I don't know how to solve it but I can produce a infinite sequence of approximations. Now I'm gonna truncate this sequence of approximations, and call this a closed form solution.
The analogy in highschool math would be computing an infinite sum that doesn't converge, but now let's instead just add to some large N, and call this a closed-form solution.
[1] e.g. https://arxiv.org/pdf/1508.06576.pdf
But how well does it fare when you give it an image of a house and an image of something completely different, like a dog or a slipper?
What is the correct answer to a question that's not well-formed?
E.g. I think most humans would say taking this content picture:
https://wallpapershome.com/images/pages/pic_hs/10150.jpg
and styling it with this picture:
https://c2.staticflickr.com/4/3499/3876547311_c2e32759d9_z.j...
is a pretty well-posed operation. How does that look using this algorithm?
I guess transfer of the wooden house amidst yellow fields with a reddening sky might lead to a wooden crab on a yellow field in front of a reddish-yellow ocean with red sky and clouds, or something.
Did you have to do anything extra to get it working? I've set things up according to the documentation (I think), but I get dimension size errors when running it.
I'm going to submit a PR, but it took me a bit of experimentation to fix these errors, so the code is a bit messier than I'd like.
I was using the pytorch 0.1.12 installed with conda (following their USAGE.md) and it took ~30s total for the transfer.
For some reason it's taking me about 4-5 minutes for the transfer, but the code now runs and the rest of the runtime is only a few seconds.
Seems fine to me. If you want to develop something commercial you'd roll your own anyway. Nothing else is restricted by this license.
The license of GCC doesn't affect the license of your binaries.
The license of python doesn't affect the license of your software.
etc.
It's really great that NVIDIA is releasing code for their deep learning research.
Virtualbox won't help you, because you can't give proper access to the GPU for the VM guest unless you set up PCI-e passthrough and dedicate your whole GPU to the VM guest (and use your integrated graphics for the host). Not sure if this is even possible if Windows is the host.
If you don't feel like setting up a Linux install on your box, you could try some of the GPU cloud services.
http://www.eevblog.com/forum/chat/hacking-nvidia-cards-into-...
Apparently this can also be done from software
Also, these posts are from 2008 and 2013, 5 and 10 years old. These hacks probably don't work any more.
This requires dedicating the whole GPU and the PCI-e slot to the virtual machine guest.
For more flexible virtualization setups, you need the professional quality cards.
A few months ago, there was TensorFire [0] that was able to do it in the browser. Quick google also gives other results [1]. There's also many apps that can do it in seconds. Speed definitely isn't an issue anymore, but getting it to work in browser can be tricky.
[0] https://tenso.rs/demos/fast-neural-style/
[1] https://reiinakano.github.io/fast-style-transfer-deeplearnjs...
Images and speech require different architectures (CNNs vs RNNs).
> conda install pytorch torchvision cuda90 -y -c pytorch
What is conda? How do i install it on ubuntu 16.014?