A Way to Detect Bias
paulgraham.com
paulgraham.com
Graham is mostly right, but slightly incorrect. In particular, suppose group A has the distribution f(x) and B has the distribution g(x).
If f(x) and g(x) are shaped significantly differently past the cutoff, then mean(H(x-C)f(x)) and mean(H(x-c)g(x)) might not agree even though there is no bias by construction. (Here H(x) is a step function and C the cutoff).
However, there is an easy fix: compute the minima of the support of the distribution rather than mean. min(H(x-C)f(x)) = min(H(x-C)g(x)) = C.
In practice, measure the weakest male and weakest female to be accepted in your sample set, or some similar approximation.
I'm pretty sure this is a valid frequentist hypothesis test. I've got half a proof worked out on paper already. It depends very weakly (and non-parametrically) on f(x) and g(x), but it works in basically the exact way Graham wants it to. Every counterexample I can think of is really pathological. My next blog post will probably be a proof of this.
All this negativity is really an overreaction. I know it's fun to totally debunk someone on details, but these are mostly fixable details.
Maybe a mean of the lower quartile would remove some the the noise.
Here epsilon is how close a typical candidate will be to the cutoff, and is mainly a function of the sample size. I don't know the behavior of epsilon off the top of my head.
This in my eyes amounts to a simplification that is neat mathematically, but removes most of the useful information. I admit though that if the minimums did not match very well, that is an obvious sign of bias.
Edit: I suppose what I'm essentially saying is, what if minimums do match, then we cannot rule out bias.
However, your claim that you can have "a token member at C for every group" is merely the claim that adversarial sampling can destroy any statistical procedure. So what? This is boring mathematically. Nor is it relevant to the problem at hand unless you want to claim that First Round Capital is actively conspiring to both a) be biased and earn less money and b) waste more money hiding that bias by funding tokens.
In practice, where do you get these definitive unbiased measurements of the expected return of investing in an individual (or graduation probability or some other score)?
It's pretty easy to measure actual return (or actual graduation rate, or actual GPA, etc.).
The test for bias proposed by pg and modified by yummyfajitas.
> before-the-fact EV
Why are you bringing in before-the-fact EV at all? That's not a component.
The test is based on comparing the minimums, yes. So, for example, what is GPA of the worst female and male students. What part of that requires expectations?
With students there is more of a connection between before-the-fact EV and outcomes, and much more information exposed (grades and course registrations on a semester-by-semester basis). Without that information (if you just looked at who graduated, who didn't) you can't really say as much. (You could run a Netflix-like competition to see who has the best graduation prediction engine and use its estimates as your definitive answer to what an unbiased EV predictor would say about college applicants.) With GPA information, let's say you look at the 10th percentile GPA's in each set and below, for boys it's 0.1-1.1 and for girls it's 0.1-1.5. What does that tell you about before-the-fact EV cutoffs? It doesn't tell you much, because one set of students, when they fail, could be more likely to fail harder than the other. It takes a lot more work than just that.
The point is that when you decide which company you finance (the source of bias) you make an estimation of future potential. But your test (and PG's one) measure ex-post results. Since there's a lot of uncertainty between the ex-ante and the ex-post measure, your test doesn't work.
Let me put in another way. The measure you are able to perform is not H(x-C)f(x) but a * H(x-C)f(x) + (1-a) * random where random is a random number and a is the weight of f(x) on the final outcome. You are right if a is 1, you and pg are wrong if a is 0 (but it will be a problem for the VCs), for everything in between you have to make assumption on the distribution of random.
I also suspect that first rounds data is useless. PG is being too credible about it - I suspect he is just being political and making his post more PC, since he doesnt seem that statistically naive. (Taking a lesson from Sam Altman most likely.)
Suppose we have to input groups, A and B. Members of each group are distributed as Exp(1), ie same underlying distribution. Our selection procedure is totally fair as well: we take everyone without question.
However, there are 9x as many people in group A as group B. So the min of group A accepted (= the min of all group A) will be distributed as Exp(9 * |B|) and the min of group B accepted will be distributed as Exp(|B|).
So in expectation the min from group A will be smaller than the min from B, and indeed this happens 90% of the time. (Aren't exponentials nice?)
Of course in this case we can note that this effect is from differences in sample size and exactly correct for it. But normally we will not know how to do this correction properly because we don't know the true underlying distribution or the acceptance criteria.
Secondly, I calculate a p-value for this test here:
https://news.ycombinator.com/item?id=10484309
It's proportional to the smaller of the sample sizes. I'll have a more detailed writeup soon. Also, this test is non-parametric, so you don't actually need to know either C or the distributions f and g - all you need is a certain level of uniform regularity in f and g.
Also, there may be a bias in the true p-value, but there is none in the p-value bound based on h(d) and min(N1, N2).
For example, assume that we find that founders who won a MacArthur "genius" grant outperformed the others. Further assume that there are only a limited number of such founders, and that all available were selected. Certainly one wouldn't want to conclude in this case that there is a bias against MacArthur fellows.
That seems obvious, but it gets trickier once you have lots of factors involved. What if the group you find to outperform consists of female founders with a PhD, substantial industry experience, and red hair[1]. Can you conclude that the process is biased against females? Males with PhD's? Anyone with red hair? Generally no, unless you are willing to assume that all of your factors are causal.
Worse, you can't even assume that it's biased against people with all of the measured factors unless you also assume that all unmeasured factors are randomly distributed. If it turns out that "a positive mental attitude" is an unmeasured but defining characteristic of success, if the interviewers rejected applicants who had less of this but you failed to include this in your category, you would be wrong to conclude that there is an unfair bias.
[1] http://www.nature.com/nature/journal/v453/n7194/full/453562a...
Alternatively, there may be a selection process in society that means only the most motivated women become entrepeneurs and so beat the average male entrepeneur.
[1]. http://www.pnas.org/content/early/2015/04/08/1418878112.abst...
That's not what the study says, it says groups that have at least 1 female founder, not female founders.
Assumption: There is no fundamental difference between a female and a male founder for achieving start-up success (average rates and variance/distribution of rates is the same)
Observation: VC funded start-ups with female founders are (on average) 60% more successful than start-ups with male founders
Hypothesis: VC funding is biased against female founders. The ones that do receive funding are better vetted, less risky, and have higher individual qualities.
Experiment: Start funding more female founders.
If we then observe: The numbers start to even out, then there is no fundamental difference. VC funding bias may have been the cause of the difference in success rate.
If we then observe: The numbers stay the same, then there is a fundamental difference and our assumption is flawed.
Rational choice: Start funding more female founders. This either removes a bias (levels the playing field), or increases your profit (funding more potentially successful founders).
PG should of course not use an hypothesis to prove an assumption (experiment/probing is needed for verification). But also: The possibility of an uneven distribution should not invalidate such an experiment (or PG's line of reasoning), it will merely bring it to light (the numbers would stay the same, thus we have shown that the difference is fundamental and not caused by a sampling bias).
An attempt at translating to mathematics (feel free to correct me!):
X = event that person belongs to group x
Y = event that person belongs to group y
S = event that person is selected
W = event that person will perform like a 'winner'
for simplicity P(X) + P(Y) = 1
Naturally, 'unbiased' in this case is simply P(S|X) = P(S1), and P(S|Y) = P(S2), i.e. that the selection process is independent of a certain variable X or Y
PG says we can measure the the performance of these selected applicant winners for each class, i.e. P(X|S,W).
I believe PG assumes that:
P(X|W) / P(Y|W) should equal P(X|S,W)/P(Y|S,W). We can see that these are different distributions, since the second is already conditioned on the selection process.
Simplified, PG assumes that P(X|S,W) = P(X|W) i.e. that conditioning on the selection process does not bias the winning results.
Its left for the reader exercise to determine the 'pathological' cases where this selection variable's distribution makes PG's assumption correct or incorrect.
However, this is simply theoretical - the actual distribution may or may not be 'pathological' and the assumptions made by PG could very well be good.
The first is if you have multiple people accepting applicants and some of them are biased to the point of not accepting applicants of particular types. That means all the applicants that are discriminated against that did make it were simply selected by people who weren't biased, and therefore won't outperform anyone.
The second is if the actual selection process is somewhat random instead of being based on pure performance. The ones who make it through that process won't necessarily perform any better, they'll just be luckier.
The third is if the application process accepts everyone equally, and then randomly prunes out people according to a bias. This is similar to the second except the acceptance criteria is still performance-based, but because it randomly throws out people (instead of throwing out low-performers), the remaining people are still going to perform the same as those who were not pruned.
The first footnote on the page also points out that if the selection criteria are different for the different groups then this process won't work, which seems like a pretty important caveat that I wish was in the article proper. One really common form of bias (especially in tech) is being biased against women, and that's also a situation where it's very common to (unconsciously or otherwise) use appearance in judging female applicants but ignore appearance for male applicants.
If noise is present, then you get convolve(H(x-C)f(x), k(x)) instead, where k(x) is the pdf of the noise distribution.
You'll need more samples to measure this, but it's completely measurable via Graham's method.
Oh. See, the problem is that if an application process is biased, and applicants perceive that bias, then those against whom it is biased will be dissuaded from applying unless they far exceed the required standards. Whereas those towards whom the process is biased will be more likely to apply, even if they are marginally qualified, because they expect to benefit from the bias.
So that means that if you do have a biased process, there's a good chance it doesn't meet criterion c - applicants in the different groups between which its bias discriminates are not equal in ability. So your test might verify a lack of bias, when there is in fact bias present.
You can't verify a lack of bias just by looking at the outcomes of successful applicants - you need to look at the outcomes for unsuccessful applicants too, to determine whether your applicant pools really do meet criterion c. Or you could look at the outcomes for nonapplicants, but that's clearly a much harder problem.
Well, they also said in their study:
> And we are not claiming that our data is representative of the industry...or even statistically significant.
Also, the wording is "startups with a female founder", not exclusively female founders... I think this is a detail that shouldn't be ignored.
And, the study doesn't show how many companies out of the 300 had female founders! Maybe it was just 1! They also say "Solo Founders do Much Worse Than Teams", so this is an important detail if there are no solo female teams ever backed in their firm! etc etc, the list goes on. Not exactly strong evidence to support the point PG is making, that bias would be easy to detect.
Measuring performance purely in terms of "how much money I make" is one way of doing it, but not the only way. And it wont cover the majority of jobs on the planet (how do you measure performance of someone who stacks shelves in a supermarket?)
I lack the mathematics to prove this, but it seems that on the face of it, pg is simply wrong. Or I'm misreading terribly.
Tangentially: Speaking of bias, why doesn't YC publish information on their companies' tech choices? PG racked up a lot of inferred cachet (positive) by stating that use of Lisp gave them a huge advantage. Now that YC has data, they should be able to show how choice of technology correlates to performance.
If that's the case, First Round Capital could profitably benefit from encouraging more female founders to apply.
Can this be true?
They excluded Uber from the results. Which, if included, makes the male-run companies look "oversuccessful". What would happen if I excluded the top female-run business, I'd bet that makes the differences between the two groups much smaller.
Given both the small sample size as well as the outsized influence of outliers, drawing conclusions from this population group is going to be fraught with issues.
When a member of a group is primed with a stereotype that their group underperforms at a task, they are more likely to underperform. So there could be a selection process biased against a group, and a selected member could be an above-average performer otherwise but, because of work environment, be underperforming.
Some universities work to remedy this through support groups or other practices aimed at under-represented minorities, and they appear to help students be more successful academically. On the other hand, there's the Hawthorne effect... https://en.wikipedia.org/wiki/Hawthorne_effect
Do you have evidence that stereotype threat hurts test performance more than real performance?
(Of course, in a hypothetical world which only eliminated the stereotype, the group would cease to be inferior. I.e., the inferiority is based on context, and is not intrinsic.)
I'll put it another way: Graham says we can look at just the performance by group to detect bias in selection. But there could be bias in selection, and A) different treatment after selection, which would not be revealed through Graham's test. Or could also be bias in selection and B) continuous reminders of stereotypes triggering stereotype threat, and this would also not be revealed.
Now you could point to a tech company with few underrepresented minorities, let's say 1%. If there's no overt bias in the work environment then people should succeed and if they don't they're just worse performers. On the other hand, if you're in the 1%, just noticing the underrepresentation among your coworkers might be a constant reminder of stereotypes.
I don't claim to have a simple solution for this.
Stereotype threat, however, is NOT such a different treatment. Again - if stereotype threat reduces measured performance by X and actual performance by Y, then the bias it introduces is Y-X. If Y=X then there is no bias.
Do you believe Y != X? If so, why?
See http://fatml.org
"For Bourdieu, cultural capital is the status culture of a society's elite insofar as that group has embedded it in social institutions, so that it is widely and stably understood to be prestigious. Schools take it as a sign of native academic ability but do not themselves impart it, performing acts of social alchemy that transform class privilege into individual merit."
There are many, many reasons that both sentences beginning "which means" are false that someone who is as smart as we're told Graham is should be able to come up with quite easily. It's astonishing that he made this tripe public.
Here's a gimme for each.
Say I'm selecting people to receive a prize; there are ten recipients and they're putatively chosen by [whatever]. But I don't like people with green eyes, so green-eyed candidates had better be pretty pleasing to me. But they can please me in any way, not necessarily in ways relevant to the metric for which the prize is awarded—maybe I also like tall people so a really tall green-eyed person averages out in terms of my predilections. They aren't relevantly better.
For the second, again, the question is "better" at what? Better at getting whatever is involved in getting selected? That doesn't necessarily correlate with outperforming anyone subsequently, especially if it's a matter of startupland. (Remember that New Yorker profile of Marc Andreessen, where Sam Altman basically admitted that he didn't know what he was doing in terms of selecting what to invest in? The flipside of that is being selected by Altman for an investment.)
Even if the VCs are totally unbiased, why couldn't the startups with women outperformed the others? It could happen for a variety of reasons. Just hypothetically speaking, maybe startups-with-women have different networking connections or insight that male-only-startups don't have?
Applications by definition are supposed to be biased. If an application weren't biased then it wouldn't be an application, it would just be a lottery. And while that's a system that a lot of charter schools actually use, or a least pretend to, I think it would be a tough sell to convince venture capitalists to allocate their capital this way.
The issue, interestingly, is that there are a lot of women (or biased-against group) between the median performance and 160% performance who are being rejected.
In other words, only the best women can get accepted, which means that above-median (but not superstars) are getting rejected, while many more above-median men are getting accepted.
This has the ring of truth to me (as pg says, it's the definition of bias).
I feel like the root erroneous assumption here is that an equal amount of people of all types are interested in the same things and that the only reason any group becomes more represented than another is that the others are getting alienated or funneled out somewhere along the way. That is a completely incorrect and invalid assumption. The fact that there are a lot more non-English-speakers in janitorial work in the US (when I worked as a janitor, I was 1 of 2 English speakers on the 12-person janitorial staff) doesn't necessarily mean the janitorial manager is biased against English speakers; it means that due to external considerations, like the fact that almost all other jobs require you to speak the native language, non-English-speakers are better suited for janitorial work, and therefore people do the logical thing, apply for work that they can do, and end up comprising a larger section of the application pool.
People make decisions based on social, cultural, and physical expectations of them, and there's not anything wrong with that. By and large, women do not have an interest in computer sciencey or entrepreneurial work. It's OK if a woman does, but it's also OK to note that most women don't. There's nothing we need to fix about it. Most women don't want to do it, and there's no reason to force them.
Why do you see fewer women becoming CEOs? Because fewer women want that kind of job and fewer women are qualified for that kind of job due to the biological realities of humanity that require women to take time out for pregnancy and child-rearing (sorry denialists, I didn't invent biology and choose that only women could bear and nurse children, so don't take it up with me), and the social and cultural expectations that have formed around these biological realities. In short, the serious applicant pool includes only a very small amount of women, so only a very small number of women obtain that position.
I couldn't disagree more. If the culture is plain chauvinism ("women belong in the kitchen not the boardroom") then there's everything wrong with that. All oppression throughout history is essentially "just culture", but that justifies nothing.
Your biological reductionism is completely at odds with our best scientific understanding of contemporary gender roles, as a few minutes on wikipedia will tell you.
Why does "culture" develop? Because people are naturally evil and black-hearted? These things don't happen in a vacuum, they develop organically because they are the best way to support human and tribal propagation and prosperity. Perhaps some things can and should change, but things that are constant across nearly all successful human societies should be considered fairly well tested.
We should note that it takes a long time to see the full effects of changes to social structures and institutions, generally at least 3-4 generations. If a society is "testing" something and the society itself expires or its success is greatly diminished within 6-8 generations of implementation, the test should probably not be seen as successful.
The West will find that traditional principles that assign gender roles based on that gender's inherent advantages and disadvantages are much more useful than currently acknowledged. Forcing people to do things that they a) don't even want to do and b) aren't well-suited for is a losing proposition, no matter how much outrage you try to manufacture to justify it.
Slavery has been tried many times but the gross inequity it inflicts means that no one can operate a stable economy or social system that depends on it.
Science mostly solved that problem decades ago, if you hadn't noticed, to the extent that breast feeding is now viewed as strange or embarrassing in certain cultures.
Medical practitioners strongly emphasize breastfeeding as the ideal form of nourishment for the baby. Formula should be used as little as possible. It's cool that we have a viable alternative solution in formula, but it's still worse than natural breastfeeding.
For a bipartisan example, consider Barack Obama and Sarah Palin.
I think as long as reasonable steps are made to avoid certain obvious bias, the rest is mostly chance.
You need to know that the probability of acceptance is conditionally independent of the "type" of the applicant given the success of the applicant.
For example, consider the following hypothesis for the First Round data: women are more honest than men. A woman presenting a bad idea to a VC will be rejected whereas a man may be able to weasel his way into getting funding. This will make men have a lower success rate, and correspondingly women will have a higher success rate.
However, this isn't really the same thing as having an across-the-board hidden bias against women.
Graham says the subjects of bias "have to be better to get selected", but what is really going on is they have to be better according to the metrics of the judge which are essentially arbitrary.
https://news.ycombinator.com/item?id=10483991
Bad measurements add noise (and increase the sample size required) but they don't invalidate the bias detection procedure.
A - 30,000
A - 10,000
A - 9,000
B - 7,000
B - 5,000 # Cutoff point below this line
A - 4
B - 3
B - 2
Even though admitting the top 5 by score is perfectly fair, the applicants from group A perform better.This strikes me as a "heroic assumption", but it's true that if you make it most of the flaws in his argument go away. Add in the unspoken assumption that the groups are both are large enough that sampling variation does not matter, and I think he's probably logically correct.
On the other hand, once you make these assumptions, the rest of his argument seems unnecessary, since all you need to know is the ratio of males and females funded. If male and female founders are exchangeable, the process is biased if one group is funded more often than they are represented in the applicants.
You don't even need to look at outcome, since we've already assumed the founders are of equal ability. I think that Paul is aiming at the case where we don't know the ratio of applicants. I think his argument can be useful in this case, but only if you have already accepted his assumptions.
Yeah. That is what happened.
I would suspect the larger issue is that people are probably much worse at identifying what performance metrics for selection convert to their respective performance for success.
The effect is the same isn't it? Less chance for a student with ancestors from a particular geographic locale getting a placement.
The "high performance means there's a bias" theory only works if you assume that everyone is starting from the same social, cultural, and physical baseline. They aren't.
Maybe a better metric would be if there were no mediocre data points among a certain group; that would be more evidence (but still not necessarily good evidence) that you have to exceptional to get attention and overcome the "bias barrier", not simply that most of a certain type of performer does better than a different type.
In [14]: x = norm(0.0,1).rvs(100000)
In [15]: mean(x[where(x > 2.0)])
Out[15]: 2.3774795090391301
In [16]: y = norm(0.5,1).rvs(100000)
In [17]: mean(y[where(y > 2.0)])
Out[17]: 2.4372124830289557
I.e., a difference in the mean of 0.5 sigma corresponds to 0.06 in Graham's test statistic.Graham is a little bit off - a better place to look for bias is the bottom of the accepted distribution than at the mean.
Look back at decisions you have made under various lenses and learn about your decisions and what biases they have, so that you can avoid them or amplify them (if positive) in future.
-> If in retrospect YC finds any factor that its selected founders who turn into unicorns ($1b, $10b etc) have in common (more than its non-unicorn, also accepted founders)
-> Then by this method, could it conclude retroactively that it had been "biased" against that factor? (since it is present more than in its non-unicorns whom it had also admitted; i.e. in other words, those with the factor are more performant than "would be expected" without the bias against it?)
Or have I misunderstood?
- As pg himself says, if we assume certain statistical distributions of ability and selection rules, the inframarginality problem goes away.
- We'd also solve the inframarginality problem if we can tell roughly who the marginal applicants were. If pg could ask the VC firm, see who almost got rejected, and compare these two groups, he'd be set. pg is well-positioned to test this on the YC dataset.
Likewise, he could solve this problem if he can observe another variable that reveals who the marginal applicants likely were (for example, the startups that had the fewest co-investors).
- There's also an entire literature out there that tries to solve the problem using other ways. For example if a system follows the "KPT" sufficient conditions then the inframarginality problem also goes away.
[1] One prominent approach ... is the “outcome test,” which originated in Gary S. Becker (1957). In the context of motor vehicle searches, the outcome test is based on the following intuitive notion: if troopers are profiling minority motorists due to racial prejudice, they will search minorities even when the returns from searching them, i.e., the probabilities of successful searches against minorities, are smaller than those from searching whites. More precisely, if racial prejudice is the reason for racial profiling, then the success rate against the marginal minority motorist (i.e., the last minority motorist deemed suspicious enough to be searched) will be lower than the success rate against the marginal white motorist. (From [3])
[2] "While this idea has been well understood, it is problematic in empirical applications because researchers will never be able to directly observe search success rates against marginal motorists. This is due to the fact that we cannot identify the marginal motorist, since accomplishing this would require having complete information on all of the variables that troopers use in determining the suspicion level of motorists. Because of this omitted-variables problem, we can observe only the average success rate of searches against white and minority motorists, and not the marginal success rate. Since the equality of marginal search success rates does not imply, and is not implied by, the equality of the average search success rates, we cannot determine the relationship between the marginal search success rates of white and minority motorists by looking at average success rates. In past literature, this has been referred to as the “infra-marginality” problem. (From [3]).
[3] Anwar, Shamena, and Hanming Fang, "An Alternative Test of Racial Prejudice in Motor Vehicle Searches: Theory and Evidence." American Economic Review. (2006)
http://economics.sas.upenn.edu/~hfang/publication/racial-pro...
There's a large literature for that, e.g.,
E. L. Lehmann, Testing Statistical Hypotheses.
E. L. Lehmann, Nonparametrics: Statistical Methods Based on Ranks.
Sidney Siegel, Nonparametric Statistics for the Behavioral Sciences.
In this case, PG will be more interested in the non-parametric case, i.e., distribution-free where we make no assumptions about probability distributions.
We start an hypothesis test with an hypothesis, commonly called the null hypothesis which is an assumption that there is no effect or, in PG's case, no bias. Then with that assumption, we are able to do some probability calculations.
Then we look at the real data and calculate the probability of, say, the evidence of bias being as large as we observed. If that probability is small, say, less than 1%, then we reject the null hypothesis, that is, reject the assumption of no bias, and conclude that the null hypothesis is false and that there is bias. The role of the assumption about the sample is so that we know that the problem is bias and not something about the sample.
In hypothesis testing, about all that matters are just two numbers -- the probability of Type I error and that of Type II error. We want both probabilities to be as low as possible.
Type I Error: We reject the null hypothesis when it is true, e.g., we conclude bias when there is none.
Type II Error: We fail to reject (i.e., we accept) the null hypothesis when it is false.
When looking for bias, Type I error can be called a false alarm of bias, and Type II error can be called a missed detection of bias.
In PGs case, suppose we have 100 startups and five of those have women founders. Suppose for each of the startups we have the data from "their subsequent performance is measured".
Our null hypothesis is that the expected performance of the women is the same as that of the men.
So, let's find those two averages and take the difference, say, the average of the women less the average of the men.
PG says if this difference is positive, then there was bias, but PG has not given us any estimate of the probability of Type I error, that is, of the probability (or rate) of a false alarm.
I mean we don't want to get First Round Capital in trouble with Betty Friedan, Gloria Steinem, Marissa Mayer, Sheryl Sandberg, Hillary Clinton, Ivanka Trump, or Lady Gaga unjustly! :-).
Let's call this difference our test statistic.
So, let's find the probability of a false alarm:
So, let's put all 100 measurements in a pot, stir the pot vigorously (we can use a computer for this), pull out five numbers and average, pull out the other 95 numbers and average, take the difference in the two averages, that of the five less that of the 95, and do this, say, 1000 times. Ah, computers are cheap; let's be generous and do this 10,000 times.
For a random number, how about starting with a 32 bit integer, with appropriately long precision arithmetic multiply by 5^15, add 1, take modulo 2^47, and scale as we want?
So, we get an empirical distribution of these differences, from the five less the 95. Looking at the distribution, we see what the probability is of getting a difference as high or high or higher than our test statistic. If that probability is low, say, 1% or less, then we reject the null hypothesis of no bias and conclude bias with our estimate of probability of Type I error 1% or less.
If with the 1% we reject, then it looks like First Round has done a transgression, will get retribution from Betty, et al., and needs to seek redemption and Betty, et al., are happy to have their suspicions confirmed. Else First Round looks like the good guys, are "certified statistically fair to women", may get more deal flow from women, and Betty, et al., can be happy that First Round is so nice!
Notice that either way Betty, et al., are "happy". That's called "happy women, happy life"! Or, heads, the women win, tails they lose, and in no event is there a huge crowd of angry women in front of First Round's offices with a bonfire of lingerie screaming "bias"!
When we reject the null hypothesis, we want to know that the reason was men versus women and not something else, e.g., a biased sample. So here is where we use our assumption of independence with the same mean.
Now we have a handle on Type I error.
Here we have done a non-parametric statistical hypothesis test, i.e., have made no assumptions, except the means, about the distributions of the male/female CEO performance measurements.
And we can select our desired false alarm rate in advance and get that rate almost exactly.
For Type II error, that is more difficult.
Bottom line, what we really want is, for whatever rate of false alarms we are willing to tolerate, the lowest rate of missed detections we can get.
Can we do that? With enough more data, yup. There is a classic result due to J. Neyman (long at Berkeley) and K. Pearson (early in statistics) that shows how.
How? Regard false alarm rate as money and think of investing in SF real estate. We put our money done on the opportunities with highest expected ROI until we have spent all our money. Done. For details, an unusually general proof can follow from the Hahn decomposition from the Radon-Nikodym theorem in measure theory, e.g., Rudin, Real and Complex Analysis. Right, in the discrete case, we have a knapsack problem, known to be in NP-complete.
What we have done with our pot stirring is called resampling, and for more such look for B. Efron, long at Yale, and P. Diaconis, once at Harvard, now long at Stanford.
Tom, with a reputation as a hacker, likes to work late, say, till 2 AM. So, we look at the intrusion alerts each minute between 2 AM and 3 AM (something like the performance of the women) and compare with those of the other minutes of 24 hours (like the performance of the men) much as above and ask if Tom is trying to hack the servers.
Or, we have a server farm and/or a network, and we want to detect problems never seen before, e.g., zero day problems. So, we have no data at all on the problems we are trying to detect because we have never seen any of those before.
So, to do a good job, let's pick some system we want to monitor and for that system, get data on, say, each of 10 variables at, say, 20 times a second. Now what?
Our work with bias in women venture applications used just one number for our measurement and test statistic. So we were uni-dimensional. Here we have 10 numbers and need to be multi-dimensional.
Well, in principle we should be able to do much better (pair of Type I and Type II error rates) with 10 numbers than just one. The usual ways will require us to have, with our null hypothesis, the probability distribution of the 10 numbers, but can only get something like that from smoking funny stuff -- not even big data is that big.
So, we want to need no assumptions about distribution, that is, be distribution-free.
So, we want some statistical a hypothesis test that is both multi-dimensional and distribution free.
Can we do that? Yup.
"You mean you can select false alarm rate in advance and get that rate essentially exactly, as in PG's bias example?" Yup.
"Could that be used in a real server farm or network to detect zero day problems -- security, performance, hard/software failures, system management errors?" Yup -- just what it was invented for.
"Attempted credit card fraud?" Ah, once a guy in an audience thought so!
How? Ah, sadly there is no more room in this post!
What else might we do with hypothesis tests? Well, look around at, right, big data or just small data.
Do we have a case of big data analytics or artificial intelligence (AI)?
Ah, I've given a sweetheart outline of statistical hypothesis testing, and now you are suggesting some things really low grade? Where did I go wrong to deserve such an insult?
Replace
Type I Error: We reject the null hypothesis when it is true, e.g., we conclude bias when there is none.
Type II Error: We fail to reject (i.e., we accept) the null hypothesis when it is false.
with
Type I Error: We reject the null hypothesis when it is true; e.g., we conclude bias when there is none.
Type II Error: We fail to reject (i.e., we accept) the null hypothesis when it is false; e.g., we conclude there is no bias when there is.
Replace
For a random number, how about starting with a 32 bit integer, with appropriately long precision arithmetic multiply by 5^15, add 1, take modulo 2^47, and scale as we want?
with
For a random number, how about starting with a 32 bit integer, with appropriately long precision arithmetic multiply by 5^15, add 1, take modulo 2^47, take the resulting integer, scale as we want for stirring our pot, and use that integer as the start of another random number?
Replace
Else First Round looks like the good guys, are "certified statistically fair to women", may get more deal flow from women, and Betty, et al., can be happy that First Round is so nice!
with
Else First Round looks like the good guys, are statistically certified fair to women, may get more deal flow from women, and Betty, et al., can be happy that First Round is so nice!
Replace
Notice that either way Betty, et al., are "happy". That's called "happy women, happy life"! Or, heads, the women win, tails they lose, and in no event is there a huge crowd of angry women in front of First Round's offices with a bonfire of lingerie screaming "bias"!
with
Notice that either way Betty, et al., are "happy". That's called "happy women, happy life"! Or, heads, the women win, tails First Round loses, and in no event is there a huge crowd of angry women in front of First Round's offices with a bonfire of lingerie screaming "bias"!
Replace
So, we want some statistical a hypothesis test that is both multi-dimensional and distribution free.
with
So, we want a statistical hypothesis test that is both multi-dimensional and distribution free.
Replace
We put our money down on the opportunities with highest expected ROI until we have spent all our money.
with
We put our money down on the opportunities with highest expected ROI until we have spent all our money. Done.
Consider two groups of candidates for a scholarship, A and B. We want to select all candidates that have an 80% or better chance of graduation. Group A comes from a population where the chance of graduation is distributed uniformly from 0% to 100% and group B is from one where the chance is distributed uniformly from 10% to 90%, with the same average but less variation in group B.
Now suppose that we select without bias or inaccuracy all the applicants that have an 80% or better chance of graduation. That means we select a subset of A with a range of 80% to 100% and a subset of B with a range from 80% to 90%. The average graduation rate of scholarship winners from group A will be 90% and that from group B will be 85%.
But we haven't been biased against A. We've selected according to the exact same perfect evaluation process and criterion from both groups. It was just their prior distribution that was different.
The actual applicant groups for jobs or financing in the real world, when they are divided by demographic factors like age, sex, race, and educational level, will almost always manifest different variances in success levels even when the averages are the same. That makes this test useless and mathematically illiterate.
And when we use a normal distribution, as we should always expect given the central limit theorem, the mathematical problems get even more intense.
This short comment is not up to pg's usual high standards for his essays.
Gee, I never saw a definition. Not sure the meaning of the phrase is clear without a definition.
As I noted in a different comment here, you can pretty easily fix Graham's test. Compute min(accepted a) and min(accepted B) instead of the means. In your example, the min of the accepted distributions would both work out to be 80%.
Suppose the cutoff sample is distributed according to f(x)H(x-C). Then the probability of the minima of a sample exceeding C+e by random chance, assuming the null hypothesis, is p = (1-\int_C^{C+e}f(x) dx)^N.
So now you have a frequentist hypothesis test. If you make reasonable assumptions on f(x) (non-vanishing near C, quantified somehow), it's even nice and non-parametric.
I.e., for any d, there is a finite probability of finding an A or a B in [C,C+d]. I don't actually care what the shapes of f or g are at all beyond this - as long as this probability exists and is bounded below (in whatever class of functions f and g might be drawn from), it's all fine.
I am assuming we know exactly one thing about the class the measures f and g come from: for every function in that class, \int_C^{C+d} f(x) dx >= h(d) for some monotonic function h(d).
The p-value is then computed in terms of h(d), since p >= h(d)^N.
Dude, your comments are normally smarter than this. Yeah, you can easily fix Grahams's test -- all you need are some numbers that do not exist and that we cannot measure.
We're talking about VC's evaluating founders. That does not, and cannot, get reduced to a numerical score. And even if VC's did use some sort of scoring rubric, then we would still not know if there was unfairness in the way they made the scores, or unfairness in the selection process. It would just be punting the problem down a layer. PG's central claim -- that a third-party can detect the bias/unfairness in the funding process just using math -- is false.
You can only know if the process is biased/unfair if you have deep qualitative understanding of the process.
That is fine, that is what he was saying. The point is that his solution is completely impractical for the original goal of finding an objective, statistically valid way of measuring whether bias exists. "Borderline" cannot be measured objectively, only by subjective rubric scoring. And when you only measure the borderline candidates, you have reduced an already way-to-small sample even further.
I made no claims about practicality - right now all I have is a little bit of measure theory showing that pg's algo is, in principle, fixable. I fully agree that the first round capital data he cites is inadequate (and also wrong, due to the unjustified exclusion of uber, which they explicitly note would alter the results).
My concrete claim: PGs idea for a statistical test is solid, I can (and shortly will) prove a toy version works, and given enough work one can probably cook up a practical version for some problems.
"Your idea isn't 100% perfect right out of the gate" is a very unfair criticism. Are we supposed to nurture every idea in complete secrecy until it is perfect?
With statistics on human affairs, 99% of the hard part is not the math, it is applying that math to a complicated, heterogenous, and difficult to measure underlying phenomena. And in most cases, statistics alone will never give you a straight answer, the best they can do is supplement and confirm qualitative observations. Failing to recognize this is how you get all those unending media reports about how X is bad for your health. PG's post was at the level of one of those junk health news articles.
This idea that statistics can only confirm and supplement "qualitative observations" (I.e. my priors) is completely unscientific and anti-intellectual. If that's true, forget stats - lets just write down the one permitted belief on a piece of paper and not waste resources on science. Science is really boring when only one answer is possible.
Since when is investing in startups a science? What is anti-intellectual, what is anti-science is to use the wrong tool for the job. Human affairs are not a science in the way that physics is a science. Statistics are far, far more fraught because there are so many variables in play, phenomena are hard to quantify, each case is so heterogenous, etc. You cannot use statistics in human affairs without also having a very good observational understanding of what is actually going on, otherwise you will end up in all sorts of trouble.
Another reason the use of mins here is not helpful, is that adding one equally awful accepted candidate to group A and B would then remove whatever bias there was according to the test, which is not what we want the test to indicate.
The idea by PG is a rough rule of thumb and breaks down trivially - suppose VC fund X were to accept all candidates, but group A was worse than B, the test would falsely imply that the fund was biased.
It's unfortunate the idea was dressed up in statistical persiflage because it isn't rigorous -- it's a rough guideline. To make it rigorous wold be very hard: either the abilities of the candidate populations would have to be measured very closely (unrealistic), or a more scientific experiment conducted (A-B test where candidates from each group are included or excluded opposite to the prior decision, which would need big groups).
Mins will fail if you conspire to cheat the test, it's true. Very few statistical tests stand up to conspiracy theories.
Lucky you. He was great. I never had the chance to take a class from him but it would have been worth a Chicago winter to have the chance.
(I took this very class the year prior)
We can illustrate it by modifing the GP example to not include time at all, and make the metric perfect.
Suppose we are selecting for the next qualifying round for the Olympic 400m team, and select candidates if their 400m is under 50 seconds. Then, we measure performance -- immediately -- by a 400m run. We have two candidate pools: people who compete in the professional circuit, and everybody else. 100% of the former group who apply qualify, while only 10% of the others do.
Okay, so now we immediately measure the average 400m time of all professionals, versus average time of all amateurs who can beat 50 seconds. It's pretty reasonable to expect that the professionals might average closer to 45 seconds, while the other group might average around 48 (I'm not a 400m expert, the actual numbers might be off. WR is 43 seconds).
According to the article, we now conclude that our selection process is actually biased _against_ professionals! This is at the least very counter-intuitive. Maybe we provide both groups with coaching, and re-test after 6 months, a year, whatever. The professionals will certainly still outperform the amateurs. However, suppose one of the amateurs goes on to greatness, and wins. Wouldn't this obviously be biased against amateurs, according to our intuition?
First, the sample size not significant. Adding back one data point, Uber, which was a real data point that was intentionally removed, likely reverses the effect.
But imagine we had real sample of thousands of companies, and it did show the result.
A typical scenario is that different demographics might connect with First Round via different deal flow channels. For instance, one channel might be longstanding personal connections, another channel might be outreach to companies in the news.
Now imagine female founders are much more likely to be found via outreach rather than personal connections. Perhaps this is due to a negative personal bias -- the VC's are less likely to be chummy with females because of their sex. So they only find female founders when their company is in the news.
It is typical in all businesses that different deal-flow channels have different average returns. So:
* If both channels perform equally well, no bias will be seen in the statistics, even though the VC's are in fact biased.
* If the outbound channel generally performs worse, then women founders in the sample will perform worse than average, even though the VC's are actively biased against them (they are ignoring all the women who would have done well, if only that had personally known them. Sine the VC's never invest in them, their results are not measured). This is the opposite of the statistical relationship that PG claims should exist.
* If the outbound channel generally performs better, then women in the sample will be better than the average.
I should also add that the differences in the channels might be due to a positive bias on the part of the VC -- perhaps they do more aggressive outbound outreach in order to get more female founders in the pipeline. Or the difference might be due to something completely neutral.
The lesson here is that using statistics is a perilous endeavor. If you want to detect something like bias, you cannot use numbers alone, you need to combine any numbers with a deeper understanding of the selection process. There is no way that a third party can run a simple correlation and determine with any degree of certainty that the field is in fact biased or not.
> Want to know if the selection process was biased against some type of applicant? Check whether they outperform the others. This is not just a heuristic for detecting bias. It's what bias means.
Under that definition, you have been biased against A. [edit: on reflection I see this as a weakness of his definition. I missed that your selection process does in fact select the best candidates.]
When asserting biases, you must first distinguish them from random noise. Using pg's logic, every selection process that isn't perfect is biased.
> Under that definition
That's not a definition. It's a claim about what the term "bias" means.
So, what are the effects of variance in different evaluation contexts, and do we have a meaningful way to measure bias if we take variance into account?
My initial reactions:
- It seems the higher the variance, the better examples you can trot out to say you are not biased against that particular group. since youll always find a member of that group who does amazing.
- If distributions of performance are multimodal, its even harder to conclude stuff because different institutions might cut off different modes when selecting the bar.
- Modeling the sources of variance may lead to insight into any actual bias.
When people complain about bias, they are not really talking about mathematical bias, but about something else: Their idea of fairness. They are talking about discrimination. And when we are discussing that, we can't really think about whether rules are applied fairly or not, but whether the rules produce the outcomes that we want.
Let's go for a ludicrous example: We'll accept all applicants whose IQ is higher than their weight in pounds. We'll be explicitly discriminating against heavy people, but at the same time, we have pretty clear implicit biases against men, and ethnic groups who tend to be taller. We might as well have said that we prefer children and Japanese women. There's no need for mathematical bias: The bias comes from the rule selection.
So, in your example, if our actual objective is to graduate an even amount of people from groups A and B, we have to, explicitly, make it easier for group B to get the scholarship. And many times organizations have objectives like that.
As a more real example, let's consider a police department. If the objective is to have a racial makeup that represents the community, and different races have different drop-out rates, the candidate selection will prefer one kind over the other, precisely to counter the drop-out differential.
So when regular people, and not mathematicians, discuss bias, the mathematical definition is unimportant. The one important thing is our stated objectives.
Alan: I believe in equality of opportunity, not equality of outcome.
Bob: How do you know there isn't equality of opportunity?
Alan: Well, just look at how unequal the outcomes are!
At this point, Bob would be wise to change the subject, because if he pressed on, he might get this: Bob: Can you give me an example?
Alan: Group X is underrepresented in Field Y.
Bob: Maybe Group X isn't as good at Field Y.
Alan: What? That's racist and/or sexist!
And if Bob were to make a comment to this effect on Hacker News, he's probably get downvoted. This is because most people agree with Alan, and many of them abuse their downvote privileges to punish ideas they disagree with rather than those that don't further the discussion. This degrades the quality of discourse, but at least it helps reassure the downvoters that they aren't racist and/or sexist.That's a pretty view, but it's inconsistent with reality, and radical egalitarians need to come up with increasingly implausible explanations to explain everyday circumstances that make perfect sense once you drop the blank slate model.
Or both. They're not mutually exclusive. Do Jamaican sprinters excel because they grow up around other sprinters or because they are blessed with natural ability? Yes. Simply put, or != xor.
A. Interviewers prefer candidates who are like themselves, interviewers are mostly white men, therefore most hires are white men.
B. The uterus and melanin both inhibit programming ability, interviewers are perfect judges of programming ability, therefore most hires are white men.
To look at the present (incomplete) evidence and decide that B is the more likely story, is racism/sexism.
a) the action of natural selection, sexual selection, and the hormone environment magically stop at the blood-brain barrier, or
b) there are real group differences between human populations?
We've already eliminated all overt discrimination. If you continue to cry discrimination, you're essentially postulating a giant unconscious conspiracy. I find the idea wildly implausible. It's much simple to just accept that not everyone is equal in aptitude and ability.
If biases affect how a professional musician hears music, is it so shocking to think unconscious bias might affect someone's judgment a candidate based on multiple fuzzy factors like ability, culture, and personality?
And that's just for job applications. You really think the criminal justice system has removed unconscious bias?
So of course the bosses pick out their friends and cronies. And a decent polity should restrain their corruption with blind auditions and accountable audits of prosecutions.
But investors should be looking for a good return on their money. They should be looking for the best investments they can find. If they're not, that is the source of bias right there.
Of course, the Wall Street industry is located in New York because you can use big city lights, strippers, and steaks to scam small town municipal pension fund managers who aren't investing their own money. Sand Hill Road is supposed to operate on different principles.
Have you ever actually worked for an orchestra? I have. The Chicago Symphony, Boston Symphony and other top orchestras take quality very seriously.
Do you think, say, Georg Solti or Daniel Barenboim were happy with "just pretty good" musicians? Their reputations (and fortunes, for top conductors are very well paid) depend on consistently outstanding performances.
And I don't know how you call it non-market. When you're income depends on millionaires donating vast sums of money, you damn well better care about quality.
It's like saying a football coach doesn't benefit if his team drafts the best players.
Let me introduce you to the extensive scientific literature on implicit bias: http://www.aas.org/cswa/unconsciousbias.html
I don't agree.
One doesn't need to advocate accusing and manhandling to think the police should use statistically valid inferences in the name of justice. I myself am a member of a minority group—men—that is responsible for a vastly disproportionate share of crime, especially violent crime. You could mandate that cops ignore this reality and treat men and women with equal suspicion, but the result would be worse policing. For example, if you look at the statistics for New York's supposedly racist "stop-and-frisk" policy, you'll find that the disparity between whites and blacks is smaller than the disparity between men and women—indeed, smaller even than the disparity between white men and black women. Why have you never heard stop-and-frisk described as "sexist"?
This inconsistency is best explained politically: complaining about racial injustice against blacks is an effective route to power; complaining about gender injustice against men is not. It's the same reason you hear constant complaints about how white tech is, but not about how black sports are. Jesse Jackson can effectively shake down Apple and Intel [1], but there is no white equivalent shaking down the NFL. (Can you imagine if "increasing diversity in the NFL" meant "increasing the relative proportion of white players"? It would be a different world—not, incidentally, one I would particularly want to live in.)
Being male means people will infer based on a superficial assessment that I'm more likely to be a criminal than, say, my sister. But that inference is correct. Being a member of such a group is my lot in life, and complaining doesn't change what is.
[1]: See, e.g., http://www.mercurynews.com/census/ci_29048321/q-jesse-jackso...
Serious question: can you at least steel man this point of view rather than making it a ridiculous straw man? If you cannot steel man it, what makes you so sure you really understand the argument?
For bonus points, you can also point out the glaringly obvious complication to this chain of logic: A. Interviewers prefer candidates who are like themselves, interviewers are mostly white men, therefore most hires are white men.
[1]: See, for example, The Blank Slate by Steven Pinker. Then, once you get over your knee-jerk "That's racist!!!" reflex, take a look—I mean actually read for comprehension—The Bell Curve by Herrnstein and Murray. Maybe add a little Cavalli-Sforza (via Steve Sailer) to the mix (http://www.vdare.com/articles/052400-cavalli-sforzas-ink-clo...). You can then graduate to basically anything by Arthur Jensen. As a topper, read "Rational"Wiki's entry on Human Biodiversity (http://rationalwiki.org/wiki/Human_biodiversity) and cringe at the smug, supercilious tone, endless strawmanning and distortion, and at the realization that you, too, were once taken in by the ridiculous "mainstream" views. (I certainly was.)
"Don`t believe any of this. It`s merely a politically-correct smoke screen that Cavalli-Sforza regularly pumps out to keep his life`s work — distinguishing the races of mankind and compiling their genealogies — from being defunded by the leftist mystagogues at Stanford."
"As you can imagine, this finding could get him in a bit of hot water if the campus thought police ever found out about it."
If you want a book length treatment, Michael Hart's Understanding Human History is the complete opposite of the typical, Jared Diamond, environmentalist accounts of human society. It is worth perusing - https://lesacreduprintemps19.files.wordpress.com/2012/11/har...
nl: Sure.
To overcome possible biases in hiring, most orchestras revised their audition policies in the 1970s and 1980s. A major change involved the use of blind' auditions with a screen' to conceal the identity of the candidate from the jury. Female musicians in the top five symphony orchestras in the United States were less than 5% of all players in 1970 but are 25% today. We ask whether women were more likely to be advanced and/or hired with the use of blind' auditions. Using data from actual auditions in an individual fixed-effects framework, we find that the screen increases by 50% the probability a woman will be advanced out of certain preliminary rounds.[1]
Bob: What? But that doesn't count because...
[1] http://gap.hks.harvard.edu/orchestrating-impartiality-impact...
http://static.themetapicture.com/media/funny-equality-justic...
The left hand side is fair rules, the right hand side shows a fair outcome
Inequality of outcome is the entire reason we see baseball played at a high level. When you demand equality of outcome regardless of talent or effort, you're asking for society to stagnate. You're asking for pervasive mediocrity. You're asking for us to kill effort and motivation. No thanks.
Kurt Vonnegut on the subject: http://www.tnellen.com/cybereng/harrison.html
Don't forget that it's these "best" people who add a vastly disproportionate amount of value to the world. They're the ones who invent new technology and discover new science. We all benefit greatly from their success.
Which essentially no one does? People favor fair outcome because they don't think the rules are, or can be, fair and often as a proxy for rules becoming more fair.
The "fair" rules would probably be to let people do stupid things and accept their own consequences, but Western culture is not willing to let houses burn down because people didn't buy into the local fire department co-op.
Anyway, I think pg's whole argument is rather moot because the three assumptions that he states are incredibly difficult to measure (Part of the reason why it is very difficult to argue for or against affirmative actions without coming across as "biased").
> We can't really think about whether rules are applied fairly or not, but whether the rules produce the outcomes that we want.
This is a more explicit way of phrasing an attitude that I've noticed in my community (a liberal U.S. university). However, I don't think it's obvious that this is the right principle to uphold.
I squirm with discomfort at the idea that we will only support "fairness" and empirical data to the extent that it is applicable to the outcome that we personally desire. This seems to imply that all evaluation metrics are "biased", until we can find a measure that selects equal representation across all demographics, regardless of the size of applicant pool or ability distribution among that pool.
What outcomes, exactly, do we want? More representation of under-represented groups? How does this relate to the goal of maximizing return on the portfolio? What does this mean for people who want a "meritocracy" (if such a thing can exist)?
Thoughts?
You can find the same inherent bias in many walks of life. Many of the hurtles to becoming a Doctor have nothing to do with being a good Doctor there just there to ensure the right kinds of people get into and out of the program.
Condemning that inequality is different from affirming that selection should be altered to produce the outcomes people view as fair. One is saying, "Don't discriminate against Xs." The other is saying, "Not only can't you discriminate against them, you need to ensure that Xs have outcome Y. That is, you may be required to discriminate in their favor."
The latter is a value, and your point about mathematics being irrelevant stands. But the former is a mathematical claim, and pg was making a mathematical claim, so the mathematical argument you replied to is relevant.
There's a genetic basis for this, as well: women have two copies of each chromosome, whereas men have X and Y, so there's no second copy to take over in men, leading to more extreme outcomes, whether good or bad.
Now of course we ought to treat every group of people fairly, but we do need to examine our priors when doing so, especially when proposing ways to detect and punish people who may be thinking bad things, consciously or otherwise.
[1] We may not know just what 'IQ' is, but we do know that tests of mental ability all correlate with each other, suggesting an underlying factor. This, in turn, can be correlated with many other things, like success (or lack thereof).
E.g. the underlying First Round's analysis likely has no statistical significance. Assuming the power law distribution of outcomes top 5 outcomes will account for 97% of value. So we now have a study with n=5.
To make the point let's apply this to YC's own portfolio. Assuming Dropbox, AirBnb and Stripe represent 75% of its value, we'll learn that YC is incredibly biased against:
* MIT graduates
* brother founders
* founding teams that do not have female founders
* and especially males named Drew
Hard to believe these conclusions are correct or actionableSee my post
https://news.ycombinator.com/item?id=10484602
where are distribution-free. So "power law", Gaussian, anything else, doesn't matter.
"C" the applicants you're looking at have roughly
equal distribution of ability.
makes the reasoning more tautological/weak.If we take two dart boards (one for female -, one for male founders) as a visual, where hitting near the bull's eye counts as "startup success".
If we take "C" to be true, then the darts would be thrown at random.
Now we draw a circle around the bull's eye. Anything landing in this circle we fund. If this circle has a smaller radius on the female dartboard, than on the male dartboard, then evidently the smaller female circle will contain more darts closer to the target (better average performance) than the larger radius male circle.
But then we do not even need performance numbers: Smaller radius circles will have less darts in them. Using "C" we only need to know that the male-female accept ratio is not 50%-50% for us to have found a bias.
In short: If you see a roughly equal distribution of ability, and (for simplicity) a roughly equal number of female to male fundraisers, then you should always have a roughly equal distribution of female to male founders in your portfolio, performance be damned.
The technique is still useful for when you do not have these female vs. male accept ratio's, and a VC publishes only success rates, but this information on ratio's is often more public than success rates/estimates.
The issue with founder funding is there are fewer female applicants than male applicants, and the applications aren't published.
I feel that a different number of darts is salvageable for this logic, but having thought about this blog post some more, I feel bias is inherently non-compute-able. Our decision on how to compute influences our results.
What PG did for me was show that there is no Pascal's wager in statistics: All outcomes/data/measurements/views are equally likely. The view that the female variable alone is able to divide skill/start-up success is weak. The assumption of non-uniform points is weak. The assumption of no variance/unequal rankings is weak. The assumption that a non-random sample is significant is weak. The assumption that VC's are unbiased in their selection procedure is weak. The assumption that nature/environment favors skilled women is weak. The assumption that decisions of who to fund does not influence future applicants. The assumption that women are still selected for capability is weak. The assumption that women ignore nature/environment and keep focusing on start-up capability is weak. It is much more likely that any other thing happens. PG's alternative is certainly a sane one, but one of many.
Perhaps women perform better because, while VC offers the same chance to men and women, they are better at picking capable women than capable men. Bias in favor of capable women.
Perhaps women perform better because, they are naturally better than men.
Perhaps women perform better because, VC is biased against women, and only the strong survive.
Perhaps women perform better because, affirmative actions to remove the inequality in performance (perceived bias) actually increased our objective bias.
Perhaps women perform better because, VC is bad at picking capable women, so they pick incapable women, of which there happen to be a lot more.
Perhaps women perform better because, now the smart and capable women start to act like the mediocre ones (bad funding decisions influence actors looking for reward)
Perhaps women perform better because, nature is "biased" against older risk-averse, but available, men and, older, unavailable women who have children, and nature favors both young males (who have to compete with the old males) and females (who compete only among themselves).
Perhaps women perform better because, our sampling method was biased.
Perhaps women perform better because, our measurements were 5 years old and we are seeing an old static state of a highly complex dynamic system.
Perhaps women perform better because, they are more variant. The good ones are really good and the bad ones are really bad, making it easier on VC's to pick the cream of the crop.
All I know is how little I know. That (algorithmic) bias is an important subject, worth thinking about, and that we need very smart people working on this subject. I would never have gotten away with upvotes on my posts in this thread if the subject was cryptography. I clearly know very little about both subjects (and only now I know that, which I hope is at least a start).
PG showed that we (I), perhaps too easily, go along with the status quo: Our measurements are all correct, our conclusions are all correct. While, if you think about it.. objectively I agree that women and men are equal in capability. If you believe this to be so, then you may have a selection bias, if you observe that men and women perform differently.
I think the least all views could do is to make sure the environment for female founders to flourish is healthy and in line with skill/capability. Then let nature do its thing.
P.S.: If we know that females actually perform better than males, what is the ethical thing to do? Fund even more female founders and make it harder for men? It would make you richer. Affirmative action? It would not remove a bias, it would introduce one.
For details, see my post
For example, if a VC funds all male founders but flips a coin to decide whether to fund each female founder, the test would fail to detect overwhelming bias.
Obviously that specific scenario is not realistic, but I believe something like this is plausible enough: A VC funds all male founders who are considered promising, and all female founders who are considered promising AND went to school with one of the partners.
And it's not hard to imagine a plausible scenario in which the test would give false positives rather than false negatives.
This essay was based on these two lines from here: http://10years.firstround.com/#one
> That’s why were so excited to learn that our investments in companies with at least one female founder were meaningfully outperforming our investments in all-male teams. Indeed, companies with a female founder performed 63% better than our investments with all-male founding teams
The comparison is not clear, but is not women versus men, but between companies with X number of males plus at least one female founder, versus those with zero female founders and Y male founders.
If we skip a step and take this fact as having some predictive value, it could be lots of things, including off-the-top-of-my-head:
1. Bias against women - which extends to teams that include men, e.g. the bias against woman exists in the presence of male co-founders.
2. That the personality traits shared by groups where women co-found startups with men are positively correlated with success. It is quite possible that these groups have much better EQ, while still retaining the IQ to impress the required amount to be selected.
3. That startups with at least one female select, and I am using this term in a very stereotyped way, "not-white-male" startups. Many Unicorn startups, from Atlasssian to Dropbox, specialise in problems faced by, again for wont of a better term, "white males". Given the mantra of solving problems we have ourselves, it is possible that mixed groups choose less male subjects. As men have been the founders of the majority of startups to date, there must be a plethora of such startup ideas left untouched. One example is DIAJENG LESTARI who started https://hijup.com/, described as "A pioneering Muslim Fashion store. We sell fashion apparel especially for Muslim women ranging from clothing, hijab/headscarf, accessories, and more." Little to no chance the archetypal "white male hacker founder" has that idea.
That's three ideas off the bat, only one of which is bias. It could still be bias, but I feel that points 2 & 3 are at least good candidates for exploration. Personally, I think there must be a lot of low hanging fruit in ideas not aimed at men, and female founders seem ideally poised to have those ideas.
This is why that isn't a counter example. Right at the beginning, the article clearly states that this method is only applicable if the prior distribution is equal:
| You can use this technique whenever (a) [...], (b) [...], and (c) the groups of applicants you're looking at have roughly equal distribution of ability.
> Now suppose that we select without bias or inaccuracy all the applicants that have an 80% or better chance of graduation.
This is subtly different than selecting for the highest graduation rate possible because it's binary, you want a group with >80% chances not a group with the best chances. Imagine if instead of the distributions you had group A was composed of people with a 100% chance of graduation and B was composed of people with an 80% chance of graduation. Our process does nothing to distinguish between those people because that extra 20% chance of graduation doesn't matter.
This brings me to what I think is the fundamental problem with your criticism, it's not clear to me what it means for a group in your example to over perform. If you select a group with the goal of 80% of them graduating it doesn't make sense to call 90% of them graduating an over performance. That only makes sense if your goal up front is to maximize the graduation rate.
I think if you rerun your example but instead assume an unbiased strategy that selects for the highest graduation rate possible you'll find that pg's essay makes a lot more sense.
>the groups of applicants you're looking at have exactly equal distribution of ability.
Rather than "roughly equal".
But obviously that makes the whole thing infeasible.
I can almost here him thinking in response, "if I throw a dog a bone, I don't want to know if it tastes good or not."
It's not clear how you can assign a candidate a 90% chance of graduation. That probability must be a subjective assessment that has come from some (biased) source. In truth, an individual will either graduate or not.
In your example, you can assign 0% and 100% probabilities in group A, but you can't in group B. The most plausible mathematical explanation for that is that you collected insufficient relevant information about candidates in group B.
Except if you want to use statistics to measure bias, you need a statistically significant sample. And actually, if you are studying complex human affairs, with a hundred different variables, you need more than statistical significance, you need a sensitivity analysis. It is similar to nutrition studies. There are so many variables at play that something can always be found to increase or decrease your risk of cancer by 50%. You really only need to pay attention when statistics show an order-of-magnitude correlation, as with the link between smoking and lung cancer.
With the First Round Capital data, they excluded Uber from their calculations, because it would skew everything. If a single data point can switch your findings to be opposite, then you just have to admit that you do not have enough data to make determination one way or another. In science it is sometimes ok to exclude an outlier, since it often indicates a measurement error. But in venture capital, you make most of your money off of the Uber-like outliers. So if you are trying to study the data to be the best venture capitalist possible, throwing out outliers is not valid.
Also, the initial premise is incorrect too. You cannot measure bias by comparing average results, because the average is not the marginal. Consider PG's footnote: "Although I used female founders as an example because that is a kind of bias people often talk about, the most striking thing was the degree to which First Round undervalue founders who went to elite colleges." Does he honestly believe that First Round is biased against founders from elite colleges?
At my last company my sense was that the MIT grads were better than the average programmer. So were we biased against MIT grads? Should we have hired more MIT grads until the average performance of MIT grads overall equaled the average performance of an employee overall? Should we have done more outreach to MIT? Should the industry as a whole hired more MIT grads?
If a talent distribution has a bunch of elite, and then a steep drop-off filled with "pretenders", then you can get this type of effect without being biased.
When we got an elite MIT grad, we hired them. When we got a "pretender", someone who was trading on the name but did not put in the work, we rejected them. And yes, I personally saw MIT grads that did terrible on simple coding exercises.
So even though the average MIT grad we hired was better than the average programmer at our company, there was no way to alter our hiring process to get more MIT grads. If we hired the marginal MIT grad that we rejected, we would have been worse off. Now we could do more outreach to MIT, and we did, but that is a highly competitive process. There were diminishing marginal returns to how much outreach we can do to get more applicants.
The statistical illiteracy of PG's post is simply stunning. Imagine a YC company gets a 100% ROI from PPC ads, and a 50% ROI from banner ads. Are they biased against PPC ads? Should they buy more PPC ads? Such an analysis is ridiculous. You look at what you are spending on the marginal PPC ad, and you stop spending when the ROI on the marginal ad is at zero, regardless of what the average is. That one advertising channel has a higher ROI on average does not mean that the company is biased against that channel.
Maybe FR is. Imagine that elite college is highly predictive of success, so you prefer to pick elite college grads, all other available evidence being equal. You're biased toward elite college grads, right?
But what if elite college grads really are phenomenally more successful, and you can't see the detailed reason (high school experience, network, whatever), to the point that they are all better than all non-elite college grads. Then even selecting 90% of your pool from elites, and 10% from the rest, is biased against the actual merit of the applicants.
[these numbers are totally made up. I'm not saying elite college grads really have these characteristics.]
The trickiness is that you can't see everything when you evaluate, so you have to assign weights to the factors you have, and leverage corellations to hidden important factor.
PG's articles are generally filled with good intuitive insight. Unfortunately, statistics can be very tricky to turn into folksy wisdom. Rules of thumb like "you need 30 samples before you can say anything" that are derived from the CLT are a good example of ones that work well enough in practice, even if they obscure some underlying subtleties. This article is an example of a rule that sounds simple, but actually has so many asterisks that one would expect it to be mostly useless in practice.
If women are performing better on average, it doesn't mean that you should invest in more women necessarily. What if all the remaining candidates would have a negative mean return? If they included Uber and all of a sudden the women now underperform men, does that mean they're biased against men and they need to invest in less women?
There's just so many statistical fallacies at play here that it's a shame that Jessica, Sam, or Geoff didn't point out that maybe someone with a stats background should read the article first before publishing it.
We assume all candidate pools are homogeneous. If all members of a certain subset A of a global population P is simply better at a task than any member of another subset B of the global population, we will see that this holds true for members of our sample as well. Thus, members of A in our sample will consistently outperform members of B in our sample. Does this mean there is a bias against A? Well, yes because if there wasn't then there would be fewer members of B or perhaps no members of B in our sample based on this result alone.
However, real life is not one-dimensional. Sometimes we need to consider other factors as well.