Transforming Wikipedia into a cultural knowledge quiz
medium.com
medium.com
This feels like a problem worth solving on Wikipedia itself. It would be nice if categories could be marked as non-hierarchical, so that for a given category, you could know whether its articles could be classified under all of the ancestor categories.
https://en.wikipedia.org/wiki/Category:Eponymous_categories would be a good place to start. Probably most of those are not hierarchical.
> This feels like a problem worth solving on Wikipedia itself.
Why would that be a problem at all? "Apple" (I assume this refers to the company) makes perfect sense in the "Steve Jobs" category. If you looked up the listing for Category:Steve_Jobs you'd expect to find it there. "Fixing" this problem would just make Wikipedia worse.
The better fix would be to realize that categories aren't strictly hierarchical, and you shouldn't operate as if they are.
(I'm next to illiterate about constraint satisfaction programming. How hard is to make reasonable crossword puzzles?)
If you wanted 'Jeopardy'-style clues, that's easier.
There was just one final problem: occasionally items would pop up that were definitely NSFW, or just made you feel icky when reading the description. To make the quiz more family-friendly, I filtered out anything related to adult entertainment (quite a few porn stars in the top 10,000), as well as contemporary people notable principally for violent crime (whether as perpetrators or victims). There are just… some things you’d rather not read about while eating lunch.
I can't help but think it would be fun to take the version with NSFW content still in, or even limited to only those items.
It would be really interesting to use that to see how certain subgroups are aware of this content. E.g. Certain subreddits and 4chan...
If Wikipedia ever breaks down its pageview stats by country, so I can generate different quizzes per-country, that's the first thing I'll do!
Did you consider making it an actual quiz with options to verify if someone actually knows what they claim? (Only skimmed the article, sorry if you mentioned it)
At the end of the day, your personal result will be for however you interpreted the instructions. But population-wide results should remain comparatively valid assuming each population is comparable in terms of being conservative/liberal with what they claim to know. E.g. you can still measure the difference between 25-year-olds and 30-year-olds because they're answering it in the same way collectively, and it's the differences I'm more interested in for research, than the absolute values.
I thought about making it something like a multiple-choice quiz, but it would have made the quiz a lot slower, and therefore either a lot longer or a lot less accurate, so checkboxes seemed like the only way to go.
I thought the explanations in the "more info" box was very useful. I'd suggest keeping it open by default on the first page (closed on other pages) as it seems a critical to read that box before taking the quiz. Just a thought!
So I would say no.
Reminds me of Kenneth Clark's definition of 'Civilization' including only Europe ...
You've got the start of something here, but is it culture or people magazine?
Wikidata looked very promising, but I was worried if it would contain all the data I would wind up needing, or if it would be in the same format 2 or 5 years from now. Wikipedia is a household name and the information in it has a lot of eyeballs on it constantly, while Wikidata as a project I couldn't tell if I could be equally confident in -- so really just taking a conservative approach is the only reason.
The communities overlap between Wikipedia and Wikidata to a large extent, but are distinct.
What I remembered was a good mixture of artists and scientists from different epochs. Now they have crammed a few scientists in a dusty corner and the rest of the museum is full of people I don't recognize.
Anyways, I am interested to see what analysis you do after you get more data.
Mathematically, there's a trick where you don't even need to compute the sum item-by-item... I calculate the binomial regression which gives me the two relevant parameters, from which I can calculate the probability density function (PDF) [1] for an item of given rank. Then I just calculate the associated cumulative distribution function (CDF) with the same two parameters [2] for rank 10,000 -- and that's the final result.
[1] https://en.wikipedia.org/wiki/Probability_density_function
[2] https://en.wikipedia.org/wiki/Cumulative_distribution_functi...
I did have a good laugh at Jared Kushner being listed as an "investor".