I happen to be working on a toy machine learning project that, based on the fictional characters known by someone, predicts their approximate age. Your list is the first organic validation set that happened onto my machine!
I happen to be working on a toy machine learning project that, based on the fictional characters known by someone, predicts their approximate age. Your list is the first organic validation set that happened onto my machine!
When Friends or Firefly were on, for example, I was watching a lot of TV (like you) because at that stage of my life I was settling into a long-term relationship (early twenties) and we found those to be mutually enjoyable things to watch. Movie watching tends to fall off around the time people start having kids. So far, on my friends-and-family (kids and grandparents alike) polls, it's pretty accurate.
Thanks again for answering :) I may just post my quick and dirty hack when I'm through playing with it.
I don't know, it's just something that I'm curious about. I would love to take your test (although I am already primed by seeing the OP's list).
It will be interesting to see how this plays out with all of the great access to media history we have today!
It guessed early 70's.
It's written to guess early/mid/late subsets of ten-year ranges, and actually does have some data for kids/teens. Mostly it had fictional characters in movies, books and television but some singers snuck in there when I had my family fill out excel worksheets, heh!
Any advice on how I could collect this kind of data? A survey on ask HN? Mechanical Turk?
EDIT: the only features it uses are fictional names that one can rattle off in one sitting and the specific age to use for the label. No gender or other demographic information. So that's what I would collect, just a list of fictional names a submitter can think of in one sitting, and their age. It's basically a form of supervised topic classification training seen in other ML tutorials, but using the age as the training set topic label. I'm experimenting on enriching the data afterwards with media (book/movie/tv flags) see if that feature improves its performance, but I'm teaching a class this week and don't have any spare time to work on it.
Looking at my data, only one person had Ted Nugent. My uncle, who is now 77. I have new respect for him. A handful of Alfred E. Neuman references are in there, wildly varying ages. It really underscores the value of a comprehensive training set to see how wildly off some predictions are when there aren't enough samples. Plus, just two vectors (as in your case) do not a classification make.
It does make me wonder about the validation though... I should try validating the model with varying numbers of input names to see if there's a baseline where it's able to reliably predict age. I think that's where media would come in handy...
Here goes, and I'm not trying to make fun of anyone, and I don't belive it's true. Personally, I believe the millenial generation might be the most knowledgable in history.
That said, I've heard some of you don't know what a 45 record is, or how a rotary telephone works?
Some people claim to not to know certain dated things because it's fashionable. I once heard a young Rebublican claim to know nothing about the GOP past. The guy sitting next to her said, "I didn't live through the French Revolution, but I know what happened?" That ended the conversation.
So what the truth?
I was barred from calling my girlfriend at night in my room, and my phone was replaced with one that was broken. My parents wanted me to be able to answer the phone, but the keypad didn't work, so I couldn't dial out.
What I could do, however, was rapidly toggle the receiver hook. Turns out the pulse frequency tolerance had a pretty wide range, and I became quite skilled at quickly tapping out her number. That, and the Nintendo gamer hotline. Priorities!