You can put a robots.txt in the bucket.
Plus you cannot put a robots.txt at s3.amazonaws.com so if the url is accessed through the https://s3.amazonaws.com/.... url, the robots.txt will not work.
I do think they should change their process (making it lazy-load instead), but that's a different issue to robots.txt.
It's manually triggered to start downloading resources every hour regardless of whether someone needs them.
In that sense, any web spider is "manually triggered" as well ;-)
This is mentioned specifically in the article, in fact.
http://[bucket name].s3.amazonaws.com/robots.txt