Object detection: an overview in the age of Deep Learning
tryolabs.com
tryolabs.com
"Unfortunately, there aren’t enough datasets for object detection. Data is harder (and more expensive) to generate, companies probably don’t feel like freely giving away their investment, and universities do not have that many resources."
Even ImageNet has only 200 classes. Imagine a person with a vocabulary of exactly 200 concepts trying to describe the world. As Wittgenstein wrote, "The limits of my language mean the limits of my world" -- and the language we can teach to computers is still very limited.
Mozilla is crowd sourcing voice data, for example: https://voice.mozilla.org
When the task moves towards determining differences between un-tagged data, you're looking at more of a clustering exercise. This is a murkier machine learning task which is relatively underdeveloped. As a quick comparison, linear regression and correlations were generally worked out by Pearson back in the late 1800's, while k-means clustering was only published openly since the 1980's.
It's also usually ideal to just have one person label one image, people often miss things or mislabel things. If you want low noise data, you need to have people repeatedly labeling the same image until they reach a consensus. The more objects the more opportunity for confusion, requiring more labeling.
You have probably been prompted to use captcha-like thing for selecting portions of an image containing a street sign. That's object detection, for a single class of object.
To answer your question, yes it should be crowdsourced, yes it is being crowdsourced, but the companies that are doing it are keeping the data to themselves because of the value/expense in collecting it.
The Tatoeba project focuses on crowdsourcing translations of sentences but has also around a few thousand spoken ones in a variety of languages. Everything is available for download for free.
It's a fun exercise in style, but does not prove much.
And I used to think this was a bottleneck before I really got into deep learning.
Actually, adding categories is very easy and does not require a crazy amount of data. It does not even requires a lot of retraining as a lot of the lower layers of a typical ImageNet model can be reused for new categories. I have made classifiers using an old model (VGG16) and less than 100 images in each new category.
If a project like wikipedia for instance started to promote the goal to have 100 different pictures for each of its articles, a classifier could be trained with any arbitrary number of categories.
http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.466...
Is there an upper limit to classes? If so what imposes this limit? Is a process like I described above ideal or is there a better way to scale to 10's of thousands of classes in image recognition?
The key is in these dense layers. If you want 30x more categories you probably need at least 30x more parameters in these. And I would naively assume that the relationship is more quadratic.
The idea is that if the lower layers trained to recognize features that are useful in differentiating 1000 categories, they probably are good enough features to recognize other categories. After all we know that there are eyes detectors there, grid detectors, text detectors, that naturally emerge.
ImageNet will continue to increase its number of categories anyway, so unless you have a cluster of TITAN GPU in your basement, you may just as well wait for their results and take their networks.
We seem to have good OD that can create horizontal bounding boxes, but these bounding boxes seem to be generic estimates.
Even rectangle detection with these models can't identify the angle or skew of a rectangle in a frame (we get the same generic bounding box)
OpenCV seemed to have models awhile ago that could do this just fine.
But if you're dealing with known geometric shapes, like e.g. rectangles, you'll get better results if you use "classic" detection algorithms that are already mathematically optimal.
For example, I once had to count the number of atoms in an electron-microscope image. I simply ran a circle detector with very sensitive settings, then culling overlapping circles with lower "circleness" score. That missed a few atoms that were stacked on top of others, but still got more accurate results than the previous method, which apparently involved a poor grad student ticking them off on paper.