EDIT: Scraper running. Should be done shortly (collecting small version, original version, XML data).
Not going to post a torrent though - mostly cuz you all can figure it out anyhow. The images are all of the form:
NNNNNN01.JPG on that S3 bucket.
NNNNNN is the object ID (the internal record ID of the work of art) zero-padded to 6 digits. The object ID is the same value that appears, for instance, in the URL on their online collection:
URL for Irises: http://www.getty.edu/art/gettyguide/artObjectDetails?artobj=...
URL for the full-rez image on S3: http://gettylargeimages.s3.amazonaws.com/00094701.jpg
You figure it out :)
Source: http://docs.aws.amazon.com/AmazonS3/latest/dev/S3Torrent.htm...
I said my original comment because the idea was to have use learn the power of wget by redownloading all files, which seemed pointless to me.
2) Iterate through objectids to fetch the following: a) xml data b) thumbnails (from getty servers) c) original work on cloudfront
You're assuming the convention holds across all ~4600 objects. Never assume. Follow the links.
That would make a convenient fizzbuzz style interview question.