They do say the dataset is for pretraining. If your application needs the ability to identify humans, then your fine-tuning dataset would include humans.
This seems totally appropriate for pretraining on natural images IMO.
This seems totally appropriate for pretraining on natural images IMO.