33 Questions
github.com
github.com
Falsehoods programmers believe about gender: http://www.cscyphers.com/blog/2012/06/28/falsehoods-programm...
Falsehoods programmers believe about names: http://www.kalzumeus.com/2010/06/17/falsehoods-programmers-b...
Falsehoods programmers believe about addresses: http://www.mjt.me.uk/posts/falsehoods-programmers-believe-ab...
Falsehoods programmers believe about time: http://infiniteundo.com/post/25326999628/falsehoods-programm...
More falsehoods programmers believe about time: http://infiniteundo.com/post/25509354022/more-falsehoods-pro...
Falsehoods programmers believe about geography: http://wiesmann.codiferes.net/wordpress/?p=15187&lang=en
[1] http://www.reddit.com/r/programming/comments/1fc147/falsehoo...
Of course, in practice, you usually target your system to a narrow set of users at first.
But yeah, if you're facebook, or work with an airline booking system, for example, you will most likely hit every single item on these lists
Reminded me of the College Humor sketch about security questions: http://www.collegehumor.com/embed/6936880/security-questions...
Yes. And they will be pissed/disappointed, especially if you're preventing them of doing something (their job, for example).
Great examples on the video! What's "impossible" today is obvious tomorrow.
I think a lot is possible with this challenge. You could compress over 8 billion yes/no questions in a single yes/no question under these rules.
And if that doesn't neatly divide people, why not let the people divide themselves through unique thought?
A 33-bit hash would probably collide too much. Yet there seems no requirement to communicate your hash back to the creator of the questionnaire with your answers. It could be a 1024-bit hash of a short story like:
Hi I am blauwbilgorgel. I currently define myself as male.
My internet names are ... I lived at ... I think we are
in Time Cube. Today is Setting Orange. I declare ...
that would create a unique hash with which to uniquely identify yourself with.I was initially confused about the first one. Then I realized that the author mentioned working on HR systems and it clicked. But for most of us who aren't building HR databases, I honestly think most professional programmers don't have to think about gender in their work nearly as much as this list suggests - the biggest thing I could come up with would be for localization, where some strings may need to be tweaked for the gender of people or inanimate objects.
Sure, but if you ask "do you consider yourself classically male" and "do you consider yourself classically female", you'll get the vast majority of people, so you can still eliminate large swaths of population with either of these.
A question like, "Are you a male or a Mexican female?" gets you even closer, though.
Not XX and not XY: one in 1,666 births
Klinefelter XXY: one in 1,000 births
There are other relevant cases, not necessarily measured here, like mosaicisms and chimerism.
"Do you self-identify as female?" is a more answerable question.
Asking about the posession of an XX chromosome should separate the people almost 50:50, the rest is something else.
On the other hand, maybe there are just <1% people who don't identify them self with being male or female, so the questions would make no difference...
Huzzah for logic. But note that the question was, "Do you have 1 apple" and not, "Choose a statement: 'I have 1 apple', 'I have 2 apples'"
> Months have either 28, 29, 30, or 31 days.
Does hold true: months do have 28 to 31 days, by definition. If someone wants my Gregorian date in a different calendar system, then we must convert, but that a display issue. Otherwise, we're comparing apples and oranges, and of course all bets are off.
It's kind of like time: the hour "2 AM" never repeats itself on random politically appointed days, because I have chosen to store timestamps in UTC. If someone wants to see that timestamp in "PST" (or America/Los_Angeles), then that's a display issue that can be accommodated, but it does not suddenly violate the fact that "2 am" never repeats in UTC.
Of course, in real life, people don't know what timezone their date is stored in and things love to break randomly on DST switch-overs or leap days or the ends of leap years.
And the 'display issues' are nontrivial - it's not just converting a timestamp to a string; it has tricky consequences for UI layout and printing if those night hours matter.
I'm curious to know in what circumstance you think a UTC day has 25 hours.
> it has tricky consequences for UI layout and printing if those night hours matter.
This is very true. I don't believe Google Calendar (or any calendar app that I've seen, really) handles displaying DST weirdness well. It would be very tricky to render. I'd love to see something attempt to tackle this.
If your days are meaningful for your app in any way, then you have to have clear boundaries between days/weeks/months - if your website has a report 'downloads per day', then it doesn't mean UTC days (which for many locations would mean splitting in the middle of business hours. And it has days where the difference between 'start-of-day' and 'end-of-day' is not 24 hours, but 25 hours.
Also, no matter how you handle time storage, if you're doing any analytics, and your process is minute-dependent instead of day-dependent (say, power consumption, not purchases), then your daily totals will have ~5% jumps twice a year that you might need to adjust.
In short, you can draw categories to include or exclude as precise a number as you like, you just have to be willing to draw really, really complicated boundaries.
I like the idea of a human UUID/GUID type identifier.
I would also like to think that this is solvable using strictly biological and physical properties, sampled at birth.
Otherwise, time and culture factors would seem make it difficult to produce a static set of "apples to apples" questions and answers.
I wonder if the right maths applied to existing genetic and forensics big data sets could produce the 33 questions.
I'd imagine that genetic markers would be the best way to do to (Disclaimer: I'm no biologist and might have made completely wrong assumptions here). They're less likely to change than, say, someone's political or religious beliefs; one could get a nasty hit on their head and forget.
The thing with genetics is that they can change over time. Some gene's turn on and off. Attributes like your face and fingerprints change over time. They're not constants.
If you could identify a set of 33 lifetime constants, you'd end up with a life-long UUID. If you expanded beyond 33 bits and included genetic markers which change over time, such as gene's which flip on/off, you could end up with a point-in-time (PIT) UUID.
UUID = Constant throughout life.
PIT+UUID = UUID plus markers identifying you at a particular point in time.
A constant would be something like, do you have a Y chromosome? (there is fault in this question: XYY syndrome)Also, you'd probably need more than 33 bits. 33 will encompass all living humans today in 2013 ADE, but would have to be expanded as the living population grows, and to include the billions of deceased humans.
In the end, a "true" unique identifier, encompassing any human, would be their UUID plus a list of all PIT+UUID's they generated during their lifetime. Or in english, an entire record of their genetics from start to end:
struct LIFETIME_UUID {
void * uuid;
void * pit_uuid_TIMESTAMP1;
void * pit_uuid_TIMESTAMP2;
void * pit_uuid_TIMESTAMP3;
...
};
That should eliminate conflicts in edge cases like identical twins or cloning.Perhaps this could be a retro sci-fi a la "Brazil", with each person carrying around a punch card with his 33-bits on them. A computer error means two people are issued the same bit pattern. In a defining shot, they hold up their punch cards up against the sun and see the holes line up. Maybe an Egyptian tomb opens too!
This is a completely wrong way to approach the problem. Because the questions should all divide the population into two parts the questions should be 'matched' to each other. This approach is a bit like doing a PCA by figuring out one component, then the other, then the rest...
One way to solve this problem is to have a lot of yes/no questions (like a big Karnaugh-table), then everybody would have a long bitstring as his unique ID. Now you need to compress that bitstring -- like the minimization of the Karnaugh-table.
http://en.wikipedia.org/wiki/Karnaugh_map
-- you need to generalize this for N number of questions (which can be done), then you'd have 33 complex questions like 'is it true that (you live in NA AND you are male) OR (you live in Canada AND you are white AND ) .. and so on and on.
Reminds me of Panoptic by the EFF: https://panopticlick.eff.org/
Everyone's ID would change as time passed (if they move, if they age, if they get a sex change, etc).
The best questions for this are inherently "irrelevant", since "relevant" questions tend to be statistically linked. So, questions like "Was the second letter of your first girlfriend's middle name between A and M?" is better than "Were you younger than 20 when you had your first girlfriend?", since we can likely guess the latter based on the other statistics.
It's very unlikely every ID will be unique if only asking 33 yes/no questions. I mean, look at two twins living together -- very few questions will be able to differentiate between them.
I think it's possible to do based on a random snapshot in time, however less possible if it's meant to last a lifetime.
I also think the questions exist, but not in a manner that we'd be able to come up with on our own. As in, I believe that a program that knew every detail about every human could create 33 yes/no questions that differentiated people, however I don't believe we could do it ourselves.
I also wonder how many questions would be required to ask non-yes/no questions and get a completely unique ID for everyone. For example, questions like "weight? languages spoken? birth place?".
I don't know if there is a simple formula for the case where different questions allow different number of answers.
I can think of almost perfectly 50/50 questions based on odd/evenness of numbers (is your current weight in grams/age in hours/number of hairs odd or even) but those are completely time dependent.
26% of global people are pre-teens; and approx half of the remainder aren't dating women.
Not even going into the fact that the social concept of 'girlfriend' isn't universal, there are multimillion cultures where the relationship stages of going from strangers to family are split differently, and none of those stages can be equated with western 'boyfriend/girlfriend' relationship; e.g. you may like someone but not be dating (so no girlfriend) and then move directly to engagement or marriage; or many other possibilities.
That means to come up with them is identical to finding an optimal compression of identifying data.
Necessarily, as the second question already implies, for this question to correctly divide the population in half, you would have to group large amounts of small populations together, resulting in very long questions.
For example, if you'd like to make another geographical question that's independent of the second one, it would have to divide in half every population of the 6 countries you mentioned. The next question would necessarily have to divide those 12 again.
By the way, the first question you ask is already suboptimal when combined with the second question, as those countries together probably do not have a clean 50% male/female split. (if they do, you should really explain that as it's not obvious)
Here is a moderately-functional question 33, though:
"Are you further North than any other person who has given the same answers as you for the first 32 questions?"
What if one of the persons was not on the planet at the time?
- Birthday (19~ bits)
- Rough Location (remaining bits)
And base the questions around those two, for example, where you born on a 1-15th, does the city you were born in start with the letter's a-k. This part would be an exercise in statistics, I would think.
edit: And one bit for if you were the first to be born of two identical twins =p
(Most people born 100 years ago are now dead.)
There is no doubt that you can uniquely identify the world population with 33 bits, so I guess the problem is to do it with questions people can answer themself.
We can go farther. People in a coma or with certain mental impairments will fall out of the system, if knowledge of any given fact is a hard requirement.
The only way this even makes sense as a thought experiment is if you have some magic biographical oracle that helps you answer the questions.
So maybe we get to assume we have some oracle that helps us simplify the hard questions.
At that stage, it's easy. Begin with, "Assume we build a list of people sorted by time of birth (with some arbitrary tiebreakers, like proximity of birthplace to Barbados, or darkness of hair color...)."
Question 1: Are you on the top half or bottom half of this list?
Question 2: Are you on the top quarter or bottom quarter of the half?
Question 3: ...
Question 1: Are you north of an east-west line that perfectly divides the population in half?
(proceed to continue to divide each section in half.)
Of course, determining the position of those lines is damn near impossible. And as soon as someone hops on a plane, the whole thing comes crashing down.
I guess it gets tricky when people are born.
Wait a minute, just realized this whole question breaks down if we don't recycle numbers from the dead back to the newly born... and there's no guarantee that all of the dead will sufficiently resemble all of the newborns. I want my money back. :)
An optimal, though inelegant solution to that goal might look something like this:
"Is the {1..33}th bit of sha1(name : location : date of birth) 1?".
Clearly you'll have tons of collisions with that solution, as you would have with any solution using 33 independent questions.
To uniquely identify people, we'd either need to use more bits, or look very closely at the population and derive very specific questions.
Why?
If we assume the hash code assignment to one of the 2^32 people is uniformly random from a set of 2^160 codes, the odds of finding a collision are astronomically small (order of 2^-95 or so). Am I missing something?
This is the birthday paradox with instead of 365 days you have 2^33 possible answer values and instead of 23 people you have 7 billion people. I leave it as an exercise for the reader to fill in these values into one of the formulas to calculate the probability of successfully giving each person a unique 33 bit answer: http://en.wikipedia.org/wiki/Birthday_problem
Right, my bad, didn't pay attention to the problem we are trying to solve :)
With 40 or 45, we could relax that a bit and use questions that are actually meaningful. Two people who are within a few bits of each other would actually be similar in ways we care about, unlike two people who are similar because their transliterated last names both appear in the last half of the alphabet.
Doesn't this sound like perfect hashing with limited memory. I don't really think that this can be done with such memory constraints. Even now we cannot produce a perfect hash function that uses 1 bit / key. The theoretical best we can do is 1.44 bit / key. And the practical best we have done till now is 2.5 bits per key. [1]
This may just be possible without the memory constraint that is , you answer N number of questions which uniquely identify you. (where N > 48 )
[1] http://en.wikipedia.org/wiki/Perfect_hash_function#Minimal_p...
Picking discrete questions like this is equivalent to building a decision tree for humanity. This is actually something that could be approached as an engineering problem (and there are mechanisms for optimizing decision trees).
The problem still remains in the face of both the technological capabilities of decision trees, and practical implementations like Hunch.com, that decision trees are reductive and discrete. Reality is neither discrete nor reductive.
It may very well be the case that there is a set of questions that could uniquely identify humans, but the insight that could be drawn from those questions might be essentially pointless.
For example:
* Were you born in the northern hemisphere?
* Were you born on an even numbered year in the Gregorian calendar?
* Is the country of your birth governed through a representative system?
It's a little spammy nowadays, but it's had enough input that it seems pretty amazingly accurate at "guessing" what / who you are thinking of in ~ 25 questions.
To properly mimic this property with yes/no questions, you will have to come up with questions that divide the whole Earth's population equally AT EVERY NEW QUESTION. Even the most obvious one, "are you (fe)male?" is slightly biased toward men (according to wikipedia). At every question that skew your 50/50, you'll have to add another question beyond 33 to catch up with this.
Not only that: The first question must divide the population in 2; the second question must divide each of the 2 subsets produced by the first question in 2; and the third has to perfectly divide the 4 subsets produced by the second... And this independently of the order in which you make the questions.
That question has some problems, but I feel like it could serve as the basis for a more specific question that would define how an only child or middle child would answer without introducing much bias. It would also have to define how half-siblings are counted and probably some other things.
We don't have true constraints on space though; why limit to 33 bits? How could we still provide a meaningful UUID to each person?
A UUID based on time and location of birth might be more feasible than any other approach, since neither will change and it's the least likely to be ambiguous. Capturing UTC at the time of cutting or otherwise removing the umbilical cord could be one way of choosing as precise, non-debatable a timestamp as any. Adding lat/long and, say, the first byte of the UTF-8 character of the mother's name (or an aspect of the mother's UUID?) could get you the rest of the way there.
Of course, this falls over in places without access to precise timing and geolocation.
2^32 = 4.29 billions,
2^33 = 8.58 billions,
2^34 = 17 billions,
For example, "are you below the median age at this exact second?" That is not a knowable answer, and changes by the second, but it does give you an exact 50/50 split.
Repeat N times for each split and we're getting very very close.
- Were you born in the northern hemisphere or southern?
$2^{33}$ is sufficient for those alive now, but the human population is a dynamic function. Set a bit when the person dies?
TODO: don't implement zombies or ghosts at this time (YAGNI principle).
Just because two things aren't statistically linked does not mean that they will never overlap.
Are there 33 genetic markers that each has no correlation on the presence of the others?
Some hard problems :- 1. Distinguish twins 2. Using characters in names as some like Chinese use non-ascii names.
In this manner, 33 bits is enough, and the birthday problem is avoided. The page mentions all this.
Or even twins who are still babies, no work required? Some cultures wouldn't even have named them yet.
First question is "Are you male?". Made me laugh.
If an early answer states the candidate lives in the north hemisphere, there's no point in asking them if they live on a landlocked African country... or whatever much more complicated questions could arrive.
Question n. Consider the number of your birth out of all people currently alive. When you divide by 2^n and take the remainder, is it odd?
I'm instead left wondering how many extra questions (35 bits? 36 bits?) it would have to be expanded to in order to produce unique results but without having to be particularly clever in producing the questions. I bet it wouldn't take as many extra as one might be inclined to think.
7,126,462,675 people * 33 bits per person / 8 bits per byte = 29,396,658,534 bytes
World population source: http://www.census.gov/popclock/
"Do you live in China, India, The United States, Indonesia, Brazil or Pakistan?" is not good question.
I ask because questions like number of siblings, favorite movie, etc. would change over time.
This is not possible unless the categories _precisely_ bisect the group each time.
And let's say you find such a question, there is no way that question would divide half of the population.
There will be corner cases, but then so does asking if someone is male.
"do you have the mumble allele?" etc.
First 33 bits of SHA512(your 3D GPS location)
You need to ensure collisions won't happen; SHA512 does not to that. (The probability of a collision in 512 bits is near impossible; in 33 bits, much less so.)
In actuality, you already had those 10 more bits of uniqueness, and people just need to ask more questions to find out which one of the many unique people you are.
This is about identification, not uniqueness.