Yolo isn't really suitable for classification - it's an object detector.
For maximum accuracy there are high-quality well tested ResNet128 or ResNet152 implementations for PyTorch[1] and TensorFlow[2] that most people would use as a basis for classification tasks where the highest accuracy is needed. Lower quality ResNets (eg ResNet50) run a lot faster.
Note that in this post, PP-YOLO replaces the older YOLO backbone with a ResNet50 (ResNet50-vd-dcn to be precise) backbone.
EfficientNet is another good option[3], as is SE-ResNet.
For this specific task it's not entirely clear what/how the classification is supposed to work though. The source of the photos is really, really important: if they are physical photos then the scanner used is important. And things like different film have different tone, and storage of physical photos matters a lot.
[1] https://pytorch.org/hub/pytorch_vision_resnet/
[2] https://tfhub.dev/google/imagenet/resnet_v2_152/feature_vect...
[3] https://tfhub.dev/google/collections/efficientnet/1