Have you verified the author's claim that this is the only match in the whole data set?
Have you verified the author's claim that this is the only match in the whole data set?
My program isn't perfect, but I only spent a couple hours on it. I think it could be improved, and better lightning when taking the photos would help too.
However, I was curious how well it did so I compared it to the author's results. My program misclassifies approximately 80 out of 27,740 Skittles.
I searched for matches and it did indeed find the same match the author found:
(0, './skittles/images/334.jpg', './skittles/images/464.jpg')
(1, './skittles/images/103.jpg', './skittles/images/384.jpg')
(1, './skittles/images/111.jpg', './skittles/images/134.jpg')
(1, './skittles/images/118.jpg', './skittles/images/353.jpg')
(1, './skittles/images/139.jpg', './skittles/images/168.jpg')
(1, './skittles/images/152.jpg', './skittles/images/281.jpg')
(1, './skittles/images/158.jpg', './skittles/images/244.jpg')
(1, './skittles/images/198.jpg', './skittles/images/255.jpg')
(1, './skittles/images/198.jpg', './skittles/images/334.jpg')
(1, './skittles/images/198.jpg', './skittles/images/464.jpg')
(1, './skittles/images/201.jpg', './skittles/images/338.jpg')
...etc
(first row is identical, subsequent rows are one different)Since there are some errors, I did get lucky. However, if it's good enough, you could also manually verify ones that are close.
Anyway, personally I'd rather write code than manually count 27,740 Skittles, but I'm grateful the author did this because it's an interesting dataset to work with.
With hindsight (although up-front simulation bore this out as well), we see that we would have to go back and manually verify anywhere from 8% (off-by-one) to 25% (off-by-two); that's every fourth pack. Now consider how much more unpleasant that manual counting is, since the Skittles are haphazardly un-arranged in the image. In short, I'm unconvinced that an automated-- while similarly accurate-- accounting would be that much more efficient.
I'm not sure I understand why sorting is necessary.
> Now consider how much more unpleasant that manual counting is, since the Skittles are haphazardly un-arranged in the image.
It's actually very easy: my program outputs images annotated with each recognized Skittle circled with the guessed color: https://imgur.com/a/jlPWXRf You don't need to manually count, just make sure all the circles are the correct colors.
I'm also pretty confident you could improve the accuracy quite a bit by improving the lighting when taking the photos, and/or incorporating something more sophisticated like a neural network (probably trained with Skittles from several bags, separated by color+rejects, with many photos in different random arrangements)
I guess this takes me a while, and seems significantly more error-prone to me than the corresponding re-count when they are sorted. Granted, it's pretty easy when the Skittles are sorted as they are in these images. But how quickly can you scan this image: https://imgur.com/a/KpddGdH and check whether there are any errors? And how confident are you in that visual check?
You are right that this approach may be good enough to find a duplicate, which was the primary objective of the experiment. But I had hoped that this might also serve as a useful dataset for future student exercises, in probability, or even in just this sort of computer vision project... but I wanted to have accurate ground truth, so to speak. Inspecting your spreadsheet, it looks like this algorithm is still less than 95% accurate, even if we only evaluate the "clean" images with Uncounted=0.