The Getty makes nearly 88k art images free to use
openculture.com
openculture.com
And 88k images is tiny when we consider that LAION has billions of entries.
Read more here: https://www.copyright.gov/engage/visual-artists/
Chapter 300 covers copyrightability in general: https://copyright.gov/comp3/chap300/ch300-copyrightable-auth...
> 313.4(A): A work that is a mere copy of another work of authorship is not copyrightable. The Office cannot register a work that has been merely copied from another work of authorship without any additional original authorship. [...] Bridgeman Art Library, Ltd. v. Corel Corp., 36 F. Supp. 2d 191, 195 (S.D.N.Y. 1999) ("exact photographic copies of public domain works of art would not be copyrightable under United States law because they are not original").
Chapter 900 covers visual art specifically and goes into more detail on the copyrightability of photographs: https://www.copyright.gov/comp3/chap900/ch900-visual-art.pdf
> 909.3(A): [...] A photograph that is merely a "slavish copy" of a painting, drawing, or other public domain or copyrighted work is not eligible for registration. The registration specialist will refuse a claim if it is clear that the photographer merely used the camera to copy the source work without adding any creative expression to the photo. Similarly, merely scanning and digitizing existing works does not contain a sufficient amount of creativity to warrant copyright protection.
Faithful reproductions of non-copyrighted two-dimensional work are considered non-copyrightable, because nothing of artistic value is added in the process. (There is a lot of mechanical work around photos, but mechanical work doesn't enjoy copyright.)
Museums like to pretend that they hold copyrights on these photos. They do not.
(In the US).
[1] https://en.wikipedia.org/wiki/Bridgeman_Art_Library_v._Corel....
don't dismiss the value in simply making them accessible. they're providing a platform to access these works, which is great.
ADHD fail on my part! I even skimmed the article and totally didn't even parse that basic bit of context. Considering that Getty Images (originally photodisc) didn't appear until the early 90s, I'd love to see this turn into a Worldwide Wildlife Federation/Wordwide Wrestling Federation thing.
It being a museum makes it a lot more compelling though. Great job Getty!
>The sad part is that the Getty gets to do these kinds of things only because it's one of the best endowed institutions on the planet with funding to pay researchers/archivists/developers to do this work
And even within those institutions there's still more work than could feasibly be done. I worked in one of the best-funded, if not the best-funded non-governmental library systems on the planet and they've got an order of magnitude more things that need to be digitized than what they've already done, and they do a LOT already.
The compilation of that collection is obviously more considerable. Archivists at Getty have already said it is about 10 years work for ~20 people.
I wasn't talking about the software development side, and I certainly wouldn't call the tech end of this a significant part of the 'intellectual work.' Organizing and managing 88k images is something my laptop 10 years ago wouldn't balk at, serving them is something a low-tier vps could do, and supporting CRUD apps for the archivists, etc are usually trivial. I think you're massively underestimating the importance of other people's tasks in these operations. Even as a developer, the people-focused work-- e.g. working with the other teams to strategize about how tech could best suit their processes, training, design iterations, etc.-- was more work than any technical component.
What makes you say that? It's already done by so many institutions within the GLAM sector. They usually don't have the same marketing budget as Getty though. Here's[1] a glimpse into the digital heritage of the EU. That's a good starting point for some exploration.
Download link is a dead dropbox account. And this is the first thing I tried.
I entered "dog" as search term and each and every item I clicked on, that had a Download option, worked fantastically well: https://www.europeana.eu/en/search?page=1&view=grid&query=do...
Taxpayer money well spent indeed!
https://www.europeana.eu/en/item/08547/Museu_ProvidedCHO_Sta...
The image is only 600x768 pixels. Way too small to be useful. You can't even read the text in the image. The original work is 30x42 cm.
There are lots of possible obstacles to showing high resolution digitizations.
Except we do.
NDL Digital Collections: https://dl.ndl.go.jp/
NDL Image Bank: https://ndlsearch.ndl.go.jp/en/imagebank
If you view items in the image bank, they show you a preview. You'll then have to click the link that takes you to the digital collections page for that item (which shows the full uncropped image). From there you will want to scroll down to the download panel and make sure you select "high resolution". Those images are generally at least around 2k x 2k or better depending on when they were captured.
You should also be able to get the unconverted image from their api (as many of them are in varying, less common formats like JP2/JPEG2000) but I haven't been able to figure out how. If you sent someone at the library an email you could probably figure it out though.
When I was there they also had an awesome special exhibit (cave temples of dunhuang if my memory and Google skills serve), which makes me think their average special exhibits are decent.
Instead of building image generators off of images scraped from people's art without their consent, we could use openly licensed images, and intentionally push the public towards licensing more of their images for open datasets by encouraging twitter and instagram/meta to add image license options to image uploads, and running some public service campaigns on these platforms to encourage use of open licenses to help build better datasets.
At the same time the smaller image sets that would be available would encourage additional research in to sample efficiency, which I regularly hear would be a generally useful area for further research.
This approach would ensure several things: Public datasets would be available to all, not just the major institutions (assuming enough people cared about open licensed models to push the major players to use them), sample efficiency would be improved, more people would get used to licensing their images for public use, and artists who did not want their style copied by these technologies would have their rights respected.
That last point would have a HUGE effect on public opinion about AI, building trust between the public and the AI research community.
Ultimately I am in favor of a world where intellectual property restrictions are sharply curtailed, but to do that ethically we would need to put in place other systems to provide for those whose livelihood depends upon their intellectual works. And if someone created a work before AI existed, and they reserved their legal rights for their work at the time of publication, I think that should be respected such that the work would be left out of machine learning datasets.
This approach would take more work, but setting this expectation would encourage the tech industry to clarify image licenses on their platforms, ultimately promoting open culture and open image licenses.
Also, I feel like the main reason artists have a beef with AI right now is mostly because lots of their published works were used without permission to train models. I think if instead SD/Midjourney et al had used open datasets curated in the way TaylorAlexander described, there would be a lot less pushback, because everyone would know the models were trained with consent of the underlying artists responsible for the training data's existence.
There is still the concern of automation eliminating demand for work done by humans, but I have a hunch that in the long term, artists will embrace these tools in the same way that's been done with Photoshop and every other digital tool. It still might be very different, i.e. AI is much more powerful/enabling than Photoshop, but I'm not sure that'll change the outcome.
I think artists will too, but so will everyone else - including many who wouldn't have the skills to create the art they want without AI and who would have had to hire an artist. Photoshop did this is a small extent, but it is possible that AI will meet most people's needs in most situations. As someone with little ability to draw or paint I'm excited for that future personally, but I can't blame professional artists for being nervous. That their own artwork is being used to train their replacement is just rubbing salt into their wounds.
Right now, AI is putting out a lot of substandard work, and artists may find themselves employed just to fix the quirks of AI output, but I doubt they'll find that fulfilling. Eventually AI art may become so homogenized, derivative, and censored that it won't satisfy clients and their customers and if that happens demand for real artists will improve, but I think things could get really difficult for many professional artists in the meantime and I don't think many will be willing to offer up their art for AI the same way most people wouldn't offer to weave rope or sharpen axes for their executioner.
AI images can copy "composition rules," but without an understanding of what the "artist" is trying to achieve, it can only guess. And the artist themself, if untrained in art, does not know what they are trying to achieve.
If you haven't read "Drawing on the Right Side of the Brain," it may help to get across exactly what I mean. Most non-artists cannot pre-visualize the image they want to create, even if it's the scene directly in front of them.
The AI art models are still going to be trained regardless. One artist opting out, individually doesn't stop any of that.
Because those art models don't need one individual artist.
There are existing art models already out there, and new models aren't going to be stopped because one artist didn't license their art.
Allowing people to remix your stuff can lead to awesome outcomes, just not personally financially beneficial
There simply aren't enough of them. Stable Diffusion is trained on 5 billion images. That kind of scale doesn't exist in public domain artwork. This dataset of 88k images is 0.0017% of that.
Also, it's worth noting that intellectual property is a weird construct. Human artists are trained from looking at copyrighted works their entire lifetime. If you asked me to draw a cartoonistic bear I cannot guarantee you that it doesn't vaguely look like Winnie the Pooh or Baloo. I've seen those things and can't un-erase them from my head. And if you prevented me from ever seeing copyrighted works for my whole life, I might not be able to draw anything.
So why are we holding AI to a different standard?
Right, which is why work on sample-efficiency would be so valuable.
> So why are we holding AI to a different standard?
Because it is an automated computer system, not a human being. It can be held to a different standard because it is an entirely different system.
For a different project I’m looking in into using k-means to determine the dominant colors.
For example I can't download the 11k version of this [1].
Is anyone else experiencing this?
Also, the server (or the load balancer) doesn't support range download, so I cannot resume where it failed :(
Edit: Upon trying to get some others, the error shows up "Read error at byte 7127040" seems they probably either limiting, overloaded servers or having some more serious issues.
The article specifically mentions Irises by Van Gogh, but the link (https://www.getty.edu/art/collection/object/103JNH) takes you to a page where it seems like you have to ask nicely and agree to terms to use this public domain image.
A publisher can certainly impose their own restrictions before publishing something to lower liability, but that does not mean the photographs themselves are not public domain, as GP comment explicitly showed.
If you have any case law since then that proves otherwise, great, but an anecdote is not proof enough.
1. Hire a photographer to take a photo of the artwork that you want to use. You have to pay the photographer, and you still need the museum's permission to access the work in a setting where you can take a good photo suitable for publication. You probably can't just snap a photo while touring the museum and use that.
2. Use a photo that the museum has and pay them and/or agree to their terms. Maybe you could take that photo and then share it because it's technically public domain, but guess what's going to happen the next time you want to use one of their photos then?
Certainly, they are under no obligation to take the photo, deliver the photo, host the photo on a server, but once you have the photo, there is no legal mechanism from preventing you from using it. Your arguments in this thread have gone: copyright protects photos of public domain works -> well people can still sue you -> well people don't have to give you access. You have arrived at the truth: museums don't have to take photos or share them with you. They own the physical artwork and can control physical access to it. I don't think that was in dispute.
That’s why it’s actually a big deal and a Good Thing that museums are making these images available online. It’s not as simple as “they were public domain anyway”.
Worse, they've attempted to extract licensing fees from people who use such public domain images, in one unfortunate case: Carol M. Highsmith, a photographer who had donated her works to the public domain received a letter of demand from Getty Images for using her own public domain images.
https://petapixel.com/2016/11/22/1-billion-getty-images-laws...
* While Getty Images and Getty Museum are born of the same "Getty" family, they aren't connected entities, and one does not reflect upon the other.
art.nelsonenzo.com.
It's fun to explore, but annoyingly he never finished it with the Metadata. So you can never search, or get them back in the same order, nor know what your looking at. He said he redirects reddit.com to there on his work computer and he just uses it to meditate. Annoying.
not too bothered with the resolution as mere mortals today can only afford to train on downsized images anyways.
this is a rather disrespectful way to refer to the benefactor.
https://brightside.me/articles/the-story-of-the-richest-man-...
I'm trying to encourage the author, and now the people here in this thread, to show some respect to the benefactor.
The benefactor has donated billions and now his foundation is offering 100k images for free to the public domain. A nice gesture.
And the best you all can do is slander him. Not the time or the place for that.
Which one of us is right?
Donating billions is not impressive at all, particularly after one's death. A billionaire can donate 99.99% of his wealth and have zero measurable impact on his standard of living, whereas a poor person donating 10% could be the difference between whether he can feed his children or not. Formation of the Getty Foundation was not charity, nor an act of kindness, nor even as you put it "a nice gesture". It was buying a name for himself in perpetuity. Lionizing the wealthy for their philanthropy post mortem is a little gross of you.
This man refused to pay the ransom on his grandchild's kidnapping, with tragic results, not because it was against principle to pay ransoms, but because he was too cheap. He ultimately allowed his son to pay, but only loaned him the ransom money at interest.
Furthermore he is not offering 100K images. His foundation directors have done so long after his death.
This man was a moustache-twirling caricature of supervillain-level miserly evil. He could have given 100 times the amount he did to his foundation and still would be deserving of derision.