PP-YOLO Surpasses YOLOv4 – State-of-the-art object detection techniques
blog.roboflow.ai
blog.roboflow.ai
The ML community has been asking the authors of that model to rename their project[2] because they are basically stealing publicity by making it seem like the next version of YOLO, despite its performance being worse than that of YOLOv4.[3]
Roboflow has deflected this in the past by claiming they don't know if "YOLOv5" is the correct name[4], but by continuing to promote it, they are directly supporting it. In fact, I wouldn't be surprised that their claim of not being affiliated with Ultralytics to be either false or a half truth, given that all the top pages about "YOLOv5" were done by roboflow, including the first official announcement.[5]
[1] https://github.com/AlexeyAB/darknet/issues/5920
[2] https://github.com/ultralytics/yolov5/issues/2
[3] https://github.com/AlexeyAB/darknet/issues/5920#issuecomment...
If they had called it FOO-YOLO, YOLO-BLAH, YOLO++, or literally anything else, it would probably be perfectly fine.
Regardless of what you think about its name, YOLOv5 a great model for a lot of use cases. And hundreds of our customers are using it in production and are very satisfied with its performance. Just as many are using YOLOv4. And EfficientDet. And MobileNet SSD v2.
They’re tools, not sports teams. It’s kind of weird that they’ve developed fanbases.
This weird defence pretty much confirms that Ultralytics and Roboflow are related though.
The many people taking issue with “v5” because it’s not by the same author as “v4” but not with “v4” even though it’s not the same author as “v3” are the “fanbases” I was referring to.
FWIW, the YOLOv4 author noted he's not opposed to Ultralytics's project (https://i.imgur.com/G00DyrX.png) as long as model comparisons are fair.
And Redmon has shared he's happy for anyone to use the YOLO name https://twitter.com/pjreddie/status/1272618558254534657
I don’t think I’m going to convince you that we don’t have some kind of hidden agenda, but we’ll continue to provide support and information about all of the new models.
But the OP is blaming the unaffiliated blog post authors for something they are just reporting.
They didn't need to do this. Part of my conversation was "I get it, you're a startup, you have to focus on business value rather than research concerns." But they made the time, and put in the effort, and I feel compelled to at least mention that that happened.
Also, @pjreddie has said that he's "happy for anyone to keep using the YOLO name! Just try to avoid version number collisions": https://twitter.com/pjreddie/status/1272618558254534657
Anyway, as a fellow researcher, I just wanted to put in a good word for Roboflow. Their priorities seem to be in order. I've also learned some interesting things from their yolo breakdowns, e.g. that training time on the newer models is significantly lower.
Heads up but insulting critics by basically calling them weird obsessed fans is not a good PR strategy. Just saying. Personally I try to avoid companies that do that since I don't know when I may end up on the receiving end for some perceived slight.
edit: Also, odd to name it YOLOv5 presumably due to the strong brand appeal of that name, and then to go and insult people for that brand appeal.
Re the edit: we didn't name it, we just reported on it using the name that its creator chose.
> Just my opinion but I’m happy for anyone to keep using the YOLO name! Just try to avoid version number collisions....
"Avoid version number collisions" means "don't use the same version number". There is nothing in that or any other tweet to indicate he doesn't think that v5 is appropriate, and if you claim otherwise you should provide a citation.
[1] https://blog.roboflow.ai/pp-yolo-beats-yolov4-object-detecti...
[2] https://blog.roboflow.ai/a-thorough-breakdown-of-yolov4/
To me PyTorch is much more convenient than Darknet.
I like yolo because it’s a production grade object defector. It seems harder to find a production grade classifier.
One amusing but dumb idea would be to use yolo for this: train the model on “photo from 1930,” “photo from 1940,” etc, where the bounding boxes cover the entire photo. But I’m curious what the professional solution might be.
I'd recommend EfficientNet (or one of its may variations) for an off-the-shelf SOTA classifier. https://paperswithcode.com/sota/image-classification-on-imag...
You could fine-tune or train from scratch. Image classification is probably the single most-researched task with the widest variety of models available. You could select an architecture based on ease of implementation, efficiency, absolute accuracy, or any combination. Papers With Code has great lists of state-of-the-art models. Take your pick: https://paperswithcode.com/sota/image-classification-on-imag...
I am not discouraging someone from building such a model, but it would be really helpful to know the context for which such a model is being developed. If it's just a hobby investigation, it would be cool to see how "predictable" dates are from images. I could even see it being used in forensics to provide a "first guess" as to when an image occurred, and helping with triaging of evidence. However, things become deeply problematic if the result from the image is fed to people as "ground truth" simply because the model was found to be accurate on a validation dataset. I certainly wouldn't want this model to be used to determine whether a suspect is innocent / guilty, or to be used naively by museums to date photographs.
For maximum accuracy there are high-quality well tested ResNet128 or ResNet152 implementations for PyTorch[1] and TensorFlow[2] that most people would use as a basis for classification tasks where the highest accuracy is needed. Lower quality ResNets (eg ResNet50) run a lot faster.
Note that in this post, PP-YOLO replaces the older YOLO backbone with a ResNet50 (ResNet50-vd-dcn to be precise) backbone.
EfficientNet is another good option[3], as is SE-ResNet.
For this specific task it's not entirely clear what/how the classification is supposed to work though. The source of the photos is really, really important: if they are physical photos then the scanner used is important. And things like different film have different tone, and storage of physical photos matters a lot.
[1] https://pytorch.org/hub/pytorch_vision_resnet/
[2] https://tfhub.dev/google/imagenet/resnet_v2_152/feature_vect...
Plug for TinyML
There have been some attempts to combine image and text data into a hybrid model but I’m not sure how widespread it is. Ex: http://cbonnett.github.io/Insight.html
Not sure if it will do anything, just curious if it would help.
For example we do this to represent the position of specific sub-image parts we extract in an original image.
> just curious if it would help
Depends what you are trying to do.
For example normal CNNs aren't rotation invariant[1] so if you know a gravity vector it can be useful to make your image upright.
(Whilst CNN's aren't rotation invariant, it's common practice to augment training data by applying some rotation to the same image, so depending on how the CNN was training it may be fine)
[1] https://stats.stackexchange.com/questions/239076/about-cnn-k..., https://stackoverflow.com/questions/41069903/why-rotation-in...
https://www.youtube.com/watch?v=OQ5LnY21Hgc
Computer vision technology, face recognition, object detection, image segmentation... it's all being weaponized.
AI/ML frameworks should have more restrictive licenses that forbid mass surveillance.