Gender Distribution in North Korean Posters
digitalnk.com
digitalnk.com
Unfortunately, there is really no good reason for someone seriously interested in accurate information, like a researcher or journalist, to use machine learning for this particular task. Labeling a couple thousand images yourself or with friends is not that big of a task. Do it over a few evenings while watching TV and drinking beer. You could have mechanical turk workers do it for you for a few hundred dollars. In either case you will get extremely reliable information. If you use multiple judges you will have a good estimate of uncertainty for every classification. There is no way transfer learning can provide this uncertainty information.
The main advantage of this technique remains the ability to quickly label very large amounts of data on the order of hundreds of thousands of rows, or thousands of columns. For smaller data, machine learning can sometimes achieve marginal improvements in predictive performance through model complexity. However prediction in smaller data regimes is mainly useful for out-of-sample prediction. The machine learning paradigm offers limited support for measuring uncertainty for out of sample predictions, which is super important if you are a researcher.
One capability of transfer learning could be to support many many applications from one model, but I have yet to see demonstrations of this in practice. The problem is that knowing how well learning has transfered requires measuring generalizabity and so cannot be done blindly.
My comment above is not about statistical variation, it is about measurement error. Inaccuracies in models for labeling images are a form of measurement error. Having humans label a (often stratified) sample of images is required to understand the measurement error. This requirement is the main difficulty when it comes to generalizing learning from one context to another. If you use multiple independent humans to label all your images, you can understand the errors made by the humans very well.
A brief analysis of gender distribution in visual representations of everyday life in North Korea
The actual title gave me the impression this was a run of the mill, shallow complaint about sexism (with, possibly, some ugly politics thrown in to boot). It took me a while to get curious about it based on other cues that this might not be the case. This piece is interesting for reasons having nothing to do with the sexism angle. It is a rich piece about history, culture, AI and possibly other things, given that I don't have time to read all the way through it at this time.
If your country has been at war for over 60 years, I suppose your society would develop a kind-of "gender-agnostic" workforce, or perhaps a total reversal of demographics from a country who is at peace.
So yes, it's likely actually true that in rural North Korea women actually do work the land (in common with many other poor countries, especially ones with food shortages and menfolk frequently conscripted or called up for army reservist drills) but are lower priority for enlistment into for resource extraction or front line military service, because North Korea needs all the farmers it can get but (in common with many other societies) doesn't see women as particularly suited to mining or guarding the country's borders. But even if that wasn't the case, it's likely the posters would still promote those gender roles. If anything, they might be downplaying the likelihood of women being drafted.
That more or less describes South Korea too.
Who knows why the author used green and yellow - feels 'blue for a boy and pink for a girl' is outdated and sexist, or just likes green and yellow better, or they're the default from their charting program, or a chosen colour scheme for the blog. But there's no doubt (for readers in the English-speaking world and who knows how many countries) what would be clearer, to the point of hardly needing a key.
Where you see both volume, distribution, percent, as well as the trendline (white) where you can see how the distribution changes - how the total count changes, and easily understand any data point at any given spot.
I wonder if/how this will change now that Kim Jong Un has promoted his sister to more visible public roles.