An MNIST-like fashion product dataset
github.com
github.com
I see no evidence at all that this particular dataset is better than MNIST. None of the issues they themselves list with MNIST are discussed with relation to their proposed replacement.
The benchmarks they provide are entirely useless - sklearn does not claim to be a platform for computer vision models. A quick WRN model gets 96% of this dataset (h/t @ajmooch on Twitter), suggesting that it doesn't deal with the "too easy" issue.
The images clearly don't deal with the problem of lack of translation invariance.
On the downside, they don't have the same ease of understanding of hand-drawn digits, which is extremely helpful for teaching, debugging, and visualizing.
We regularly get high 90's accuracy on real world images where the system has to auto-crop all on its own. Inventory images are far too easy.
What's more, the actual hard cases (like tiny shorts vs tiny skirt, or long blouse vs short dress) would be near impossible to do at such low resolution.
There are more than a dozen image classification and segmentation examples on the scikit-learn gallery:
1. Scrape images and store as png
2. Downscale to 28px
3. Convert each image to grayscale
4. Convert to matrices and add label (additional row?)
5. Normalize to have matrices of 1 and 0 for faster computation
6. Vectorize said matrices
7. Concatenate into one big vector
Did I miss something / Am I fooling myself?
I plan on working on my first ML side project and I would love to gain some insights from HN.
1. Yes but you need to manually inspect and verify that the images are of the right class
5. Images are grayscale, not only black and white.
Additionally, MNIST and fashion-MNIST have all their objects centered and of similar scale. This is a large part of what makes them a popular first test for any image model: they are very simple to solve as the model need not be very robust to fit the dataset.
Yeah, this is crucial. Especially when trying to generalize models. It’s easy to verify the usefulness of data augmentation when you can make basic assumptions about the data (e.g. it’s centered).
5. I’d provide a vector of integers where each integer represents a different class. Encoding is dependent on the underlying algorithm and its implementation. You’d want to provide plenty of flexibility to users.
I’d also add that MNIST has an advantage over a dataset like CIFAR because background pixels are zeroed out (i.e. you don’t need to account for varied backgrounds). So you’d probably want to segment your objects.
And our research on recommenders using it: http://sharknado.eggie5.com
Particularly, the 2D scatter of the CNN features: http://sharknado.eggie5.com/tsne
But just like MNIST, it seems to lack variety in the positioning of the important elements, they are all centered which means that they don't train the network in being translation invariant. I presume this issue can be tackled with data augmentation techniques like applying affine transformations.
However, I am not sure if we need more MNIST-like datasets. With small size many things make much less sense (data augmentation, even convnets as images are centered anyway) plus using many channels is a typical things (IRL I rarely work with grayscale images). So I am curious, in which way this dataset is better than CIFAR-10?
See my note on datasets in Learning Deep Learning, http://p.migdal.pl/2017/04/30/teaching-deep-learning.html#da....
https://i.imgur.com/viV7gFB.png (x-axis: Fashion, y-axis: MNIST)