CNN-generated images are surprisingly easy to spot for now
peterwang512.github.io
peterwang512.github.io
Also, let me quote the linked page:
> malicious use of fake imagery is likely be deployed on a social media platform
Nope, that couldn't possibly relate to any news network. Never never never gonna admit that!
I mean personally I'm all in favor of more usage - or even automatic insertion - of the `<abbr>` tag. Can probably be done with a browser addon as well.
https://trends.google.com/trends/explore?geo=US&q=%2Fm%2F0x2...
The ability to classify photos by news outlet based on identifying their photojournalism rules through computer algorithms sounds like a remarkably clever idea.
HTML doesn't have that ambiguity.
I this case, however, there’s a conflict with the news network which could also plausibly be the subject of the headline. They have interenational recognizability, and have been using the acronym almost exclusively for years; it is effectively their name.
We are not computer algorithms here. A human being can decide "yeah this sounds like cable news network" and use the long form of this CNN.
The day someone uses “HTML” to mean “hyper-threaded machine learning” or whatever, yes definitely.
CNN was unambiguously used for the TV channel for decades now, of course some people are confused when one uses it to mean something else without warning.
HN is suitable here because it can be assumed that Hacker News denizens are acquainted with a rather obvious shorthand for their own community.
YC... likely as above but might be safer explicated and explained.
Perhaps my pre-caffeine morning brain is overly pedantic but Generative Nets use deconvolutions to generate images from latent codes, so using CNN rather than GAN (Generative Adversarial Network) is a bit confusing in this context.
CNNs are used by VAEs (Variational AutoEncoders, also generative) use convolutions to produce the latent codes and the discriminator (adversarial) part of GAN training uses convolutions.
I think Generative Networks ( or GNNs ;-) ) would perhaps have been clearer.
Did they really hoped for that their paper will remain in a specific group of experts? I seriously doubt so.
And the website isn't published in the CVPR, it's published on the internet.
Some cases I've seen lately seem to forgo this not out of ignorance but as a form of eletism/knowledge gate keeping.
It's a natural tendency for ingroups. Nearly any video game forum, or anything else that's full of hobbyists will ultimately contain posts that are absolutely full of acronyms. And they're impenetrable. Bear in mind, I'm not defending this behavior, and certainly not disagreeing with you.
And I'm not saying this is _my_ solution, this was literally taught to me in engineering first year.
If you're writing a paper, define every acronym the first time you use it.
If you're in a forum with a set of acronyms known to all, define them in a sticky or the forum readme.
For example: images generated by convolutional neural networks (CNN) are easy to identify.
May be this is specific to I.T or Computer Science? Where there are thousands of abbreviations and acronyms which itself is often the name people use. SQL, DRAM, CPU, HTTP, SRAM, FPGA, URL, TCP/IP, UDP, NAT, DHCP, GPL, etc.
I mean if you are discussing technicals of Neural network you expect your audience to at least know CPU, GPU, and FPGA. And if you are discussing software development I hope I dont have to spell out GPL.
So I dont think it is a form of eletism/knowledge gate keeping. In the age of internet you can search those "acronyms" meant without the full name, which isn't something could be easily done 15 to 20 years ago.
In other industry such as Mobile Wireless Networking, those acronyms are often clearly spell out because there are comparatively little of it. FDD, TDD, MIMO, NR or LTE are often spelt out in full when they first use.
It doesn't have to be a hard rule, but major topics of a subject should be spelled out, at least, then you're giving people something to work with in their web search.
I'm my university first-year CS class, a third of the students had never heard of GitHub. Now, that's easy to look up, and GPL seems to be a lucky acronym as well, but CNN certainly isn't. Expanding it the first time or adding a footnote costs you nothing, but people not right in your field or still learning tremendously. Someone who got their Master's in CS 10 years ago likely wouldn't have heard of CNNs at all, and neither would most new CS students.
A scientific publication in such a broad field with such a widely-applicable topic and one of the most clashing acronyms right in the title should most certainly at least expand their key terms.
UK reports rampant student marijuana use before class
That headline has quite a different meaning if “UK” is abbreviating “United Kingdom” versus “University of Kentucky”.
"However, these methods represent only two instances of a broader set of techniques: image synthesis via convolutional neural networks (CNNs)."
It's true that once you've trained your CNN you could make a non-convolutional NN that computes exactly the same things but less efficiently, but the point of an NN is not just what it can compute -- there are lots of systems that can, given enough parameters, approximate arbitrary functions well -- but how you train it.
Edit: Although, I see it does in first use in the introduction, so maybe that's just conforming to whoever's style guide.
Would this be receiving as much attention if they had used "Convolutional Neural Networks" instead of just CNN?
If not, this might indicate that the fingerprints or artifacts left by the generators are not of the "perceptible" variety.
Also a discriminator trained from this experiment might be useful to train a more powerful generator.
Also, if the images are not recognizable as fakes by humans, then it's good enough. What would be interesting in going further than that? I actually see it as a feature if at the same time it's possible to prove when images are fake.
(1) We show that when the correct steps are taken, classifiers are indeed robust to common operations such as JPEG compression, blurring, and resizing.
(2) When using Photoshop like methods the detector performs at chance (is useless).
(Full disclosure: I have only read the abstract so far.)
If I understand the terms used, it sounds like you're suggesting adding this classifier to the discriminator, to avoid detection. Since they are already failing to pass their existing discriminators, it seems like they could try to not be detected, but they wouldn't actually succeed.
Could the classifier that they're using here be used as a discriminator in a GAN, to help train it to avoid this detection method?
But sadly I could not obtain a clear picture of what is the difference between their detector and a baseline one. There are some minor points and references about upsampling, downsampling, resizing, cropping and fourier spectra comparison across generators, but those seems to be just comments and comparison and not crucial points in the construction of the detector. Furthermore data augmentation doesn't play a big role, they say that it usually improves (a little) the detector.
As a math person I like to get some more meat from papers, but here it seems that little tricks allow then to win the game. Perhaps that is the way (little or no math involved) to make advances. Well, at least they say that shallow methods modify the fingerprint of the fourier spectra so that now you can't detect which is the generator of the image.
Perhaps the "universal word" was what captured my attention.
Instead, I haven’t seen a proposal for a system I think could work well. Camera and phone manufacturers could have their devices cryptographically sign each photo or video taken. And that’s it. From that starting place, you can build a system on top of it to verify that the image on the site you’re reading is authentic. What am I missing that makes this an invalid approach?
I do understand that this would require manufacturers to implement, but it seems achievable to get them onboard. I even think you get one company like Apple to do this and it’s enough traction for the rest of the industry to have to follow suit.
I am coming at this from the angle of, who would use this type of service other than the courts? Certainly major news organizations could benefit but we have numerous recent examples where they have either run with CNN imagery but they have also purposefully run video and use images of similar events to portray the view they wanted for a current event.
Of course in the end, if the end game is to have news, image, and video, validation there will need to be more than one and in separate enough areas of the world to have some chance all would not be intimidated / infiltrated to the point they are not trust worthy
So:
1) train a network that can detect CNN generated images
2) train the CNN network to generate whatever you want, politicians in compromising positions, etc. but also add in weights against the the other network
3) Images won't be easy to spot...
People will obviously start writing CNNs that detect images that are generated obfuscated this way with CNNs, but still, it's all possible.
Typically; a GAN (Generative Adversarial Network) consists of (1) the generator; a model generating images and (2) the discriminator; a model that learns whether images it is fed come from the generator or from the image dataset. The (gradient) information of how the discriminator made its decision is fed back into the generator, in order to help it learn how to generate more _real_ images.
The discriminator is what you describe in step 1, and the generator is your step 2.
Next, it is possible to have a test which detects those. And that test can be improved by better training.
Then, another AI learns how to synthesize images which the fake image detector AI can spot, until it learns how to fool the fake image detector.
Then the fake image detector is improved by training it against the improved fake image synthesizer.
Repeat.