Using Waifu2x to Upscale Japanese Prints
ejohn.org
ejohn.org
Still, in terms of pure shock and awe - they're jaw-droppingly nice for upscaled versions, to the point where if you didn't have the original, it wouldn't occur to you that this wasn't it.
Re: Image 2, http://i.imgur.com/541uG5t.png
Since JPEG uses an 8x8 block transform, you can find artifacts by shifting the image a few pixels over and looking for how the transformed block changes, basically.
https://github.com/FFmpeg/FFmpeg/blob/master/libavfilter/vf_...
Also, using a better chroma upscale can help for small images. libjpeg just uses nearest-neighbor (no "real" resizing) and hardly anyone notices, but it helps with lines and edges.
That's precisely what it's attempting to smooth over - and it works well for anime images because in them, those variations would be considered noise.
Wondering if it would anime-ise it, I fed Waifu2x the standard Lenna image - twice - ended up with this:
This is trained to upscale anime images and not woodblock prints - and anime images are typically flat, uniform colors. This may have issues scaling up a still of a background scene from 5cm/second but would fair much better with a character still from .hack//sign. You have to keep in mind what it is trying to scale up.
>(Naturally I could train a new CNN to do this, but it may not even be necessary!)
Training to upscale woodblock prints and retaining the texture might make sense if you care to retain texture. It only works as well as it does because the style is very similar.
Although it's not clear what scenes exactly the upscaler was trained up, I suspect that it's currently best suited towards scenes that have lots of large bold lines and not lots of tiny details.
A common solution to resizing anime characters is to create a colored vector of an image. The differences between these vectors and the original stills are minimal and usually 'satisfactory'. There is an entire scene of people who create these vectors and another scene of people who use the vectors to create wallpapers and other graphics. [0] Waifu2x can help replace the need to vector these images by increasing the quality of upscaling them.
This is the prevalent 'style' for anime - at least from the past 8-10 years or so. There are a few outliers and I imagine Waifu2x would work poorly on them. For example, I do not see it working well on a still from "The Garden of Words". [1]
[0] http://img04.deviantart.net/96b1/i/2015/105/8/d/oumae_kumiko...
[1] https://24framesps.files.wordpress.com/2014/11/the-garden-of...
I tried it on the actual source myself (using http://waifu2x.udp.jp/), and there was very little actual loss of this kind.
(Edit: It looks like the low-noise-reduction version was added later, and you were talking about the high-noise-reduction version, in which case, fair enough.)
I tried it with just upscaling and no noise reduction, and the result is about what you'd about: a really nice upscale, perfectly preserving all those patterns (as well as the noise, unfortunately). Doing that and filtering in another program might work better.
Image 1: http://i.imgur.com/4cXr51v.png
Image 2: http://i.imgur.com/PZAXeM8.png
It doesn't come with any noise reduction, but nothing stops you from doing that separately from the upscaling process itself, and that way you should be able to control it better anyway (I find the reduction options provided by waifu2x really aggressive even with the low setting, it just kills tons of detail).
As a sidenote, when talking about something like image scaling, it would be a good idea to avoid saying something like "image scaled 2x (normally)" as there are lots of ways to scale images and what's "normal" can vary a lot depending on what you're using.
You're absolutely right that I shouldn't have said "normal". I update the post to clarify that this was using "OSX Preview". I did some hunting but didn't find any obvious pointers as to which algorithm they're using. If anyone knows offhand I'll be happy to include it!
Also chat with @deepbluecea who's done a lot of image processing stuff, including for Apple.
I have/continue to use imagemagick and similar software-based solutions and they're pretty slow for multi-MB images (but most servers don't have good GPUs so it's the only solution unless you're building custom racks as imgix does).
Especially if you set imagemagick to use the much worse scaler that imgix uses, I imagine it'd be pretty fast.
On the other hand, if you replaced imgix's stack with the high quality scalers from mpv (written as OpenGL pixel shaders), and then compared to expensive CPU scalers, I would expect a GPU solution to be a win.
Note that imgix also has to recompress the image as PNG or JPEG at the end. This has to be done on the CPU and is probably more resource intensive than any of the scaling.
[1] http://www.vapoursynth.com/
Its benefits are questionable if used on anything else.
I've also tested it on anime screenshots and in that case it's it pretty much is en par there with NNEDI3 (which is computationally much cheaper) because real world encodes actually have compression artifacts and those get scaled too if you disable noise reduction or everything is smoothed out too much if you leave it on.
So if you want to use it on anything else you really do have to retrain the NN first, otherwise you get results you could also achieve by other means (e.g. warpsharp, NNEDI or Photoshop Topaz)
Also, waifu2x only scales luma. Its chroma handling is just regular upscaling (whatever imagemagick uses by default. I think), so even that part could be improved.
[1] http://forum.doom9.org/showpost.php?p=1722990&postcount=3
http://research.microsoft.com/en-us/um/people/kopf/pixelart/
> You also wouldn't want to use this to guide a police sketch...
You would, could and can. Take a look at police sketch software, it is a very manual version of this process. There are a very limited number of potential variations of the human face, that is why eigenfaces work for facial recognition. Consider the scenario where you have a photo of a human face, where one half is occluded. In manually reconstructing the image, you wouldn't place the reconstructed eye an inch above the chin - because human faces don't work that way. Neural networks pick up on that. There is software that tailors use in fitting suits, where a few measurements (like weight and arm length) can be used to extrapolate the rest of the measurements (like chest size and torso length). This works because of the limited number of potential human dimensions.
As far as use of these techniques for evidence... I'd actually prefer it to reliance on eye witness accounts, as the algos are open to exact measurement - unlike most of the other stuff that passes for "criminal science" (humans are still in the loop for fingerprint analysis, wtf?).
As are eye witness accounts, which have been demonstrated to be pretty useless.
As are fingerprints, a tiny sliver of (maybe?[0]) uniquely identifiable information.
As are autopsies, where the state of the corpse is maintained only in whatever the examiner writes down, x-rays, or snaps a polaroid of.
As are bite marks...
So you've got all that, plus your lawyer's sweaty appeals to emotion in a group of 12 people - of whom four will express a belief in haunted houses and two will claim to have actually seen a ghost [1]. You'd prefer that over an application of math that can be challenged and rationally discussed?
> Furthermore, neural networks are trivially fooled...
A neural network was fooled with the equivalent of a hash collision, one guess as to how to fix that :)
> I'd actually trust a trained neural network far less than a human...
I can't think of a single person I'd trust over math, once maybe Bill Cosby - but not anymore.
> Speed and automation are their advantages compared to trained humans, not quality.
Well in this context I'd say that impartiality and repeatability are pretty important, which are characteristics more likely to describe a math model than an individual with the qualifications of a mailing address - and all the training that can be packing into a 20 minute vhs about civic duty played on a wheeled TV.
[0] http://www.academia.edu/447251/The_Current_Position_of_Finge...
[1] http://www.pewresearch.org/fact-tank/2013/10/30/18-of-americ...
Norman Tasfi made a Neural Net upscaler for Flipoard http://engineering.flipboard.com/2015/05/scaling-convnets/
I expect video upscaling next.
There is a directshow filter (madvr[1]) for windows that already offers a neural network scaler {NNEDI3, simpler network than waifu2x) in realtime.
"The final use case that we thought of was saving bandwidth. A smaller image could be sent to the client which would run a client side version of this model to gain a larger image."
Can be applied to gifs and videos, but this really depends on the usage case and if the client would tolerate such a thing.
Anime is almost never drawn with finer detail than the output resolution, so artifacts are not a problem. This is a low resolution scan of something with very fine detail, something which it is not trained on.
At worst it seems comparable to the previous result. At least to my eyes.
On the other hand, it does help you grok its function, and I suspect the 'memorable' name is at least partially responsible for its popularity.
If you really are working with the original source, you should rescan to png or tiff or even just higher rate jpeg?