Open Questions about Generative Adversarial Networks
distill.pub
distill.pub
Regular journal articles are published as PDF files, which make a decent electronic analog of physical paper. In the case of Distill, there's the usual DOI, volume, and issue, but the page itself isn't static. I did find a seemingly undocumented feature referenced in a GitHub issue - you can append "index.archive.html" to the url to get a supposedly standalone html archive of the page. However, when I tried this with a number of recent publications there were missing images and broken formatting all over the place, seemingly due to the authors having linked in external resources (primarily for a number of the figures).
Does anyone know how to do this or have any ideas?
It's not that I have any pressing need for offline copies at this precise instant, but they have proven useful on a number of occasions.
It's "a tool to push web resources into web archives", which basically entails making requests of archiving services (like the Wayback machine, for example). Using `archivenow --all <url>`, you can back up <url> to a whole bunch of different services, and get links to snapshots of that page at the time of access.
This is mostly sufficient, so long as at least one service is trustworthy and doesn't shut down unexpectedly. One issue that might come up is that the archives might miss content if the user has to interact with the page a bunch in order to reveal it, but that's unlikely to be an issue for the sorts of things an academic would be archiving.
As another poster mentioned, some bibliography managers have this capability baked in, or can acquire it with a plugin.
However, if you want more control, you might consider using something like `pywb`[1] to make your own web archive, and then host it yourself (e.g., on S3 or something). For more robustness, you could also publish those WARC files as torrents, or on IPFS, or if you're really cool, as part of the transaction messages on some blockchain (provided that you can stomach the transaction fees and don't anticipate having a large audience in China).
The above will probably suffice so long as the Internet continues to function as a going concern. To accommodate concerns about what happens in the case of total societal collapse, I've been backing up the papers that I cite by converting to binary[2] and preserving the result via aluminum punchcards, which I then mail to Sam Altman since he's got a way better game plan for the end times than I do[3].
I'm not sure if aluminum is optimal for this purpose though, both for material science reasons and because I've been informed that I am under investigation for my role in some sort of aluminum smuggling ring. If that fails, my last backup plan entails constructing an allegorical representation of the papers in question (which is surprisingly easy-- the Hero's Journey from Joseph Campbell[4]'s exegesis is basically about differential equations anyways), which I will then craft into a compelling mythology and teach to the local youths. Hopefully this will ensure that the citations are preserved via oral tradition even in the wake of cataclysmic change. Figuring out the BibTeX for this archival format will be pretty tricky, though.
But yeah, unless you have some extra aluminum lying around, `archivenow` or whatever's built-in to your reference manager should do the trick.
---
0. github.com/oduwsdl/archivenow
1. github.com/webrecorder/pywb
2. After stripping most of the markup removing images using `pandoc`, as aluminum is expensive.
3. https://www.bloomberg.com/features/2018-rich-new-zealand-doo...
That's going to be one HUUUUUGE mythology... I'm looking forward to it!
If you have an idea of what we could do on our side with a fixed amount of work, please let me know! Unfortunately we can't require our third-party authors to be experts in graceful degradation. I'm sure better tooling could help, but currently have no specific inspiration.
As a fallback for your specific use case—we do actually provide a tiny bit of print styles in our CSS. Try your browser's PDF conversion/print feature; I find it often produces acceptable results. (Minus proper alignment to pages, of course. :/)
---
From my perspective, the main issue seems to be broken formatting and media (alignment, hoverable citations and notes, missing or partial images, etc). The recent activation atlases paper is a particularly good example of the problem (https://distill.pub/2019/activation-atlas). I don't have time to tally up all of the differences right now, but briefly just a few big ones so we have something concrete to refer to here:
* The index.archive.html version shows "Loading..." in place of all the images in the first figure (and most of the other figures as well).
* The Zotero web page snapshot version actually does have the images for the first figure, but that figure still has other issues. Many of the other figures still just show the "Loading..." text though.
* In both cases, text is flowing around the figures in strange and unreadable ways.
* The bibliography for the index.archive.html version is completely broken.
* In both archive cases, uMatrix is showing third party scripts running on the page. I've allowed them, but obviously loading things from remote locations means it isn't really standalone anymore.
---
> we haven't been able to think of a satisfying automatic way of providing static fallbacks to interactive diagrams.
Why does it have to be static? Why worry about graceful degradation at all? I guess I just find myself wondering why there would be two versions of a scientific paper (ie "real" and "archive") in the first place? I don't see any reason why scientific work needs to be published in a static medium, but it does seem like it should always be standalone and immutable.
To that end, why not require submissions themselves to be fully standalone entities with absolutely no external dependencies (ie review them without internet access) and then just publish those? Then instead of pdf papers, you end up with interactive (but standalone) html papers that are viewable in any standards compliant web browser, scripts and all.
Barring such a drastic solution, I don't have any clever ideas from a technical standpoint; I'm not a web design expert nor am I familiar with what tooling you allow authors to make use of in their papers. A fairly blunt approach might just be to require authors to provide static images as standins for each figure. It could literally just be a static image of what the interactive figure initially looks like before you do anything with it. Then you'd just have the existing formatting issues to sort out, but presumably those are fairly straightforward.
An aside: Figures don't seem to be numbered in the activation atlases paper I referenced above. It might be nice to add a requirement to clearly number and label figures, tables, and other such insets - as far as I know, all the major academic publications do this. If you're wondering why, at minimum it makes it much easier to discuss the paper when you have a clear label such as "fig 3a" or "fig 4 panel 2" or whatever to refer to.
> Question: How does copyright work for GAN output? If I input 300,000 copyright protected photos of celebrities and generate images of new celebrities that do not exist, are the generated images public domain or would there be copyright issues?
AFAIK, this is not a settled issue, but I'd be really interested hear what an actual lawyer thinks about this?
Maybe the work is derivative of a particular training image if we can easily predict the presence or absence of that training image given only the trained model?
A human artist is a legal person, a GAN is not. That's a fairly substantial legal difference.
If another animal such as chimp draws a human face what do you think would be the outcome? Would a chimp drawing a face be any different to a human drawing a face?
Law is only extremely rarely concerned with only output results and not process, status of actors, etc.
> Would a chimp drawing a face be any different to a human drawing a face?
Yes, legally, the outcome wouldd be different (whether it was original, a direct copy, or something made by copying elements but with some new content) because chimps are neither legal actors that can create a copyright through authorship nor ones who can violate copyright.
See https://en.m.wikipedia.org/wiki/Monkey_selfie_copyright_disp...
What's so interesting with ML is that there's a dense spectrum between memorization and originality. I hope in the future people start checking how much their models are memorizing.
My favorite case is Google's facts, like when you google "golden retriever weight." From other people going out, measuring it, writing it up, and publishing that info to the web, Google can extract the info and never direct traffic to those sites. I still don't know whether I think it's okay.
Link: http://www.dmlp.org/legal-guide/works-not-covered-copyright#...
Even that has a spectrum, from "What is harry potter book is first?" to "What is the first line of harry potter?" to "What is the text of the first harry potter book?"
A lot of Google's answers to questions veer towards being written opinions and original definitions anyway
One of the things that has most amused me about creating https://www.thiswaifudoesnotexist.net is watching the reactions to images being posted elsewhere, especially without (initial) attribution: not a few people like the faces enough to request info about the character or artist, or say that it looks a bit machine-learning-like but is obviously too good to actually be ML-generated, or compliment OP on their illustration! People have begun casually using them as avatars, and there's one account on Pixiv for just uploading TWDNE or other faces generated by my StyleGAN models.
As I understand it, the cases established a multi-part test for whether a particular use of images is infringing. One of the key parts is whether the use is “transformational”—I would argue that most GAN output would fall under this fair use exception for transformational use and thus would be ok (at least here in Calif).
In the aforementioned cases, thumbnailing a la Google images was sufficiently transformational so synthesizing whole new images certainly should be. The tests do, however, depend on the use so I wouldn’t say that re-synthesizing similar-but-different images with a NN is always and automatically fair use.
I have a background as a lawyer (albeit not in IP) and currently working with ML, and I can see a case being made either way.
https://www.artnome.com/news/2019/3/27/why-is-ai-art-copyrig...
> Claims that AI is creating art on its own and that machines are somehow entitled to copyright for this art are simply naive or overblown, and they cloud real concerns about authorship disputes between humans. The introduction of machine learning as an art tool is ironically increasing human involvement, not decreasing it. Specifically, the number of people who can potentially be credited as coauthors of an artwork has skyrocketed. This is because machine learning tools are typically built on a stack of software solutions, each layer having been designed by individual persons or groups of people, all of whom are potential candidates for authorial credit.
I'll answer what I think you're asking and you can tell me if I got it wrong:
There's been a lot of effort spent on coming up with different loss functions for GANs, but https://arxiv.org/abs/1711.10337 shows that, according to the metrics in problem 5, they don't really improve results compared to the original GAN loss function.
There's something called a Wasserstein GAN https://arxiv.org/abs/1701.07875 that you may have heard about, but IMO the useful thing to take away from that paper is their 'gradient penalty' technique and not the new loss function.
Using complex NN models (let's say, LSTM) doesn't give much better results than simple neurons with deep layers. Seems that the model extraction (how to represent some author style) is the only point that matters, and it then easily forged (i.e., not hard to reproduce the desired style on purpose).
I'm trying to use GAN as a mean of desconstruction, of removal of the ease patterns to see if some NN can use subtle characteristics to identify the author.
IIUC, that's something that's been looked at in the ML Fairness literature. See this paper for example: http://www.aies-conference.com/wp-content/papers/main/AIES_2...
Related: do GANs work well as a data augmentation technique and can their exact contribution to how much a model could potentially be improved be quantified.
EDIT: Added some clarifications
Encrypting the data secures the data loading part but it's not like your model didn't learn anything.
There might be encryption schemes which make learning such a "language" harder for an rnn, but it does not necessarily mean these methods are harder to break. So it might be possible to find/design an encryption method which is easy to learn (map x to y), but hard to break. I don't know much about encryption though, so I might be wrong.
Proof: You can easily train an RNN to return the n-th letter of its input and ignore everything else. Take that function to be f_n. Then g_n is the function that takes the encrypted input and returns the n-th letter, encrypted. If the encryption scheme is learnable in the above sense, then you can also train an RNN to perform that task. You can use multiple of those RNNs to extract an encryption of each individual letter in the input, which is equivalent to having the input encrypted with a substitution cipher. Substitution ciphers are so easy to break with frequency analysis that people have been doing it by hand for more than a thousand years.
Conversely, if you have a secure encryption scheme, then all statistical regularities in the input should be obfuscated so that they cannot be exploited by an RNN.
I can think of two ways that procedure can fail:
1. If some aspects of the private data aren't represented in the public dataset (e.g. codewords in classified documents), the final classifier won't recognize them. On the other hand, I'm not sure whether there's an actual use case for making a classifier public that's required to work on data that can only be obtained from a non-public source.
2. If some aspects of your labels are sensitive (e.g. which political opinions you agree with) then labeling the public dataset retains that information. In that case, you'd likely need to apply techniques similar to those studied for removing racial or gender bias from a model.