Why Nature will not allow the use of generative AI in images and video
nature.com
nature.com
Allow me a bit of a rhetorical question, what are the chances they already publish photos taken on devices that apply by default some form of AI-based generative/corrective algorithms like the "AI detail enhancement engine" by Samsung (the one they use to enhance photos of the moon)?
Isn't this a contradiction, though? My understanding is that generative AI is created entirely from software, using a network of previously created images as input. A corrective filter modifies an image taken directly from a sensor instead.
I personally don't mind aesthetic corrective modifications to photos. I was an astronomy observatory last night and learned that most of the magnificent images we've seen of distant nebulas and galaxies have post-production coloring applied, they mostly look black and white coming off the sensor. Does the coloring fundamentally change our understanding of what it is that we're looking at? I don't think so, and that's where I draw the line.
"or augmented"
It's interesting to compare this to other situations where, say, law tries to create lines that aren't really there and the incentive to ignore imaginary ones is greater than the incentive to keep them.
This seems to be a very common phenomenon with technology.
I doubt that the editors are under some illusion that the nominal ban will create a hard line in reality. I'd be surprised to learn that that is their idea of success with this measure.
Your understanding is incorrect, generative AI can modify an image taken from any source, as well as creating from scratch.
As a result, most findings should be validated by verifying that some property of interest is present in the high dimensional raw neural data, though that's only conceptually possible sometimes.
Every photo is touched by "corrective algorithms". Nature is talking about generative AI specifically, which means using an LLM to generate part or all of an image. This precludes using Midjourney, Photoshop's new "generative fill", etc.
I assume that what Samsung's "Space Zoom" feature does — replacing elements with higher-quality stock photography — was already disallowed. If so, whether the elements were identified/replaced manually or automatically isn't really a concern from an editorial perspective.
e.g. "Digital images submitted with a manuscript for review should be minimally processed. A certain degree of image processing is acceptable for publication (and for some experiments, fields and techniques is unavoidable), but the final image must correctly represent the original data and conform to community standards. Editors may use software to screen images for manipulation."
[1] https://www.nature.com/nature-portfolio/editorial-policies/i...
Hey ChatGPT, write a prompt for midjourney to generate a realistic photo of XYZ with ABC parameters.
Then plug it into Mid-journey.Technically the LLM isn't generating the image, and I agree, but I think their point is rather obvious and we need not be intentionally obtuse nor needlessly pedantic.
There are quite a few in-depth explanations of the whole system; here's one for instance: https://jalammar.github.io/illustrated-stable-diffusion/
Also, what constitutes "raw data" is itself a matter of debate. How raw is raw? Like any interesting pursuit, scientific publishing struggles to keep up with developments in technology.
Certainly no jpeg image produced by any digital camera is really “raw” as it will already have been through a debayering filter
https://en.wikipedia.org/wiki/Bayer_filter
And then on top of that is the JPEG compression artifacts.
But I do wonder how many raw files also contain data that has been debayered already. I have not looked into that.
I know that with third party firmware such as Magic Lantern it is possible to get the image data without debayering. https://magiclantern.fm/
Likewise I know that the Camera Module 3 for Raspberry Pi is possible to retrieve the image data from without debayering.
There isn't a difference between auto white balance and generative AI though. The colors in an auto mode digital camera picture are not real.
I'm going to keep that in mind, there does seem to be this interesting human nature presumption that everyone keeps in sync with the latest and greatest. But that's simply not the case.
Where the line is drawn as to what's "generative" and what's "AI" may be blurry, but they haven't just banned traditional transform operations.
It's on the rise for the past years
100%?
The reality is we all know what kind of images to expect from Nature. Generative Ai is not appropriate there and we all know it.
eg 18 May 2023[2] "The cover shows an artist’s impression of two male mammoths fighting"
or 20 April 2023[3] which shows the DART spacecraft, apparently photographed from nearby in space.
[1] https://www.nature.com/nature/volumes
If I just want some random artwork, like an image of, I don't know, a blackboard, why is using Generative AI inappropriate?
Ironically, Nature's own licensing rigor drove me to generate this art. It was replacing content that had come from other sources, where the time to obtain and clear copyright was too long for our timeline. More hilariously, one of the images that I replaced was from the US government, and in the public domain. The other was from a consortium in which I am part of the project leadership.
They seemed perfectly okay with this, as long as I proved to them that I had the professional Midjourney account where copyright is not encumbered. I wonder when they will again allow this kind of use.
I just asked DALLE for “A scientific illustration of a membrane bound protein being phosphorylated”, and while the results aren’t all that credible, I could imagine using them as a starting point.
I wonder if their graphics designers will need to move from industry standard software to something less capable. Interestingly the Amish may have been ahead of their time in creating purposely limited technology that was compatible with their beliefs (https://www.npr.org/sections/money/2013/02/25/172886170/a-co...).
Adobe - https://www.adobe.com/sensei/generative-ai.html
Figma - https://www.figma.com/community/plugin/1145446664512862540/A...
Some new tools have popped that are centered around generative AI (I have no idea if they’re any good):
Prototypr - https://prototypr.io/toolbox/diachat
Diagram - https://diagram.com/?ref=Welcome.AI
How do you verify whether this cartoon illustration of of stacks of money against a red background is "accurate and true"?
https://media.nature.com/lw767/magazine-assets/d41586-023-01...
Would it have made a difference if that image were generated by Midjourney?
The actual reasons, given later in the article, are that Nature is taking a political/legal position on copyright and privacy. That's fine by me, but it's disappointing that they give a misleading and nonsensical justification before the actual justification, as if to make their stance sound less political.
But it is irony of the ironies for Nature which sources all its content AND revisions from the open community to say they care about fair copyright compensation of creators.
Aren't they unlegitimizing their own business model by claiming such things?
Why not just apply this rule to all media? What is the purpose of singling out images and video?
So you'd have to make X=0.0001 or something, and then what? Pay them all a fraction of a cent?
In general, it's extremely difficult to prove that anything is AI generated at all. Even more impossible to prove which model was used with which settings.
In the case of artwork, the author of even the most convincing, artifact-free AI generated piece will immediately crumble if asked to show WIPs, non-flattened project files or timelapses. I have seen some charlatans attempt to fake WIPs by using style transfer to turn their finished piece back into a "sketch" but the results aren't very convincing, the models aren't trained on the process of creating art conventionally so they're not good at faking it.
I’ll say that even in my personal life if I catch you flat-out lying to me about something I have a very difficult time reestablishing trust. It’s like you’ve revealed that deep down you think it’s an acceptable behavior and now everything that comes out of your mouth has to weighed as possible bullshit.
It seems like it could be pretty simple — if there's a question, you ask the creator to provide the original RAW and have a conversation about how they got to the final "developed" image. If there's still doubt, they could be asked to duplicate/approximate the process in a screen-sharing session.
I'm not familiar with the current state of content provenance initiatives like Content Authenticity Initiative¹, but generative AI is likely to boost their popularity.
¹ https://en.wikipedia.org/wiki/Content_Authenticity_Initiativ...
Also, is generative AI capable of dramatically upscaling the quality of it's output relative to its input? I would assume so but I've never really thought about it.
You could cheat and convert the image to a RAW file, but it'd be very difficult to do so in a way that would fool a forensics investigator.
> Also, is generative AI capable of dramatically upscaling the quality of it's output relative to its input?
If the image output is too small, one could use tools like Topaz Gigapixel AI to scale it up
Cameras have been doing that for at least a decade, but that's not foolproof either. For example, Canon's Original Decision Data was cracked in 2010. https://photographybay.com/2010/11/30/canon-original-data-se...
The reason it's hard to fake RAW files isn't because one can't convert images to RAW files, but because RAW files contain lots of additional information that would be difficult to fake. For example, RAW files include mosaiced sensor data which has flaws that are unique to a particular sensor.¹ A digital forensics expert can evaluate a hundred aspects of RAW files to see if anything smells fishy.
¹ https://www.labmanager.com/sensor-imperfections-are-perfect-...
https://www.forbes.com/sites/danielfisher/2012/01/18/sopa-me...
The publication already has a reputation and I don’t think people would judge Nature if they used Midjourney for featured images.
Videos are an entirely different thing, it will take a few more years for AI to be able to create interesting videos, so in a sense it is meaningless to even mention it.
Dictating the tools that artists use for a commission is punitive and moralizing.
Let the artists decide the morality of their own profession.
It's a straightforward clarification of their existing editorial policy. https://www.nature.com/nature-portfolio/editorial-policies/i...
I know that's a naive truth and we all know it. But still, we really do pretend otherwise.
I think that might be a bigger deal than we acknowledge. I think maybe our sanity is bent from living this way.
Evidence that STEM people can think clearly about this, when their paycheck doesn't depend on pretending otherwise.
(Personally, I'm going to be in the latest AI techbro gold rush, but will try to do it responsibly.)