If the math is wrong, show us.
If the math is wrong, show us.
However, when it comes to mammograms, where the false negative rate is up to 12.5% and the false positive rate is up to 50%, and the incidence of breast cancer in women (12.5% over a lifetime) is rather high, false negatives would play more of a role.
The math is, given the test shows a positive, what is the probability that you actually have the disease?
If we are not making that assumption, then the 50% chance claim at the end is incorrect, since there's a chance that both of the two positive results are false positives and there is a false negative somewhere else in the remaining 99 (out of 101) people.
The 50% chance of it being a false positive (conditional on two positive results being observed) is actually a lower bound, and will increase as a function of the false negative rate.
A 1% false negative rate has a negligible impact on the math. 50% isn't perfect down to the decimal point, but it's very close.
After the first round you have 99.99 false positives and .99 true positives.
After the second round, if independent, you have .9999 false positives and .9801 true positives.
And in fact the lower bound, if you had no false negatives, is slightly below 50%. .9999 vs 1.
Their math shows that those are the same number, and they didn't elaborate on definitions.
The least misleading definition of accuracy is that accuracy = 1-false_positive_rate = 1-false_negative_rate. So unless told otherwise, I will use that definition.
"Let's say there is a terrible disease that 0.01% (1 in 10,000) of the population is infected with, and we develop a test that is 99% accurate. You go to the doctor and get a test. Oh no, it's positive. What are the chances that it's a false positive?
Intuition would tell you 1% since the test is 99% accurate. Intuition, as it often is, is wrong. Take a sample size of 'x', we'll use 10,000. How many people will be infected? 1. How many false positives will there be? 1% * 10,000 = 100"
Then they conclude by saying:
"So there will be 100 false positives, and 1 real positive."
The fact there is one real positive necessarily implies they are assuming the false negative rate is exactly 0, given that the disease prevalence is in 0.01 percent of the population.Or that they rounded .99 people to 1 person.
And given that they did round the false positives, I don't see why that's a strange idea.
First, I'm assuming that top comment meant the false-positive rate was 1%, even though it was poorly stated as "accuracy." If you're commenting on this original ambiguity, then the whole thing is ambiguous.
But if we accept a 1% false-positive rate, then it makes no difference at all what the false-negative rate is, once we have a positive test.
Out of a 100,000 tests, on average 1000 will be falsely positive. That's what a 1% false-positive means. It makes no difference how many are truly positive, how many are truly negative, and how many are falsely negative. The definition of 1% false-positive means that, out of 100,000 tests, on average 1000 will show a positive test when the condition isn't actually there.
So if 100,000 people each took two tests, on average 10 will have two false-positive tests. Again, this is independent of the number of false-negatives, and is simply the definition of false-positive. (Some tests may have correlated false-positives, i.e. you're more likely to get two in a row, but that's not standard and not part of the definition.)
So if you have a disease whose incidence is 10 people out of 100,000, and you took two tests which are both positive, you are equally-likely to be in the group of 10 people who really have the disease, and the group of 10 people sitting there with two positive tests who don't have the disease.
You seem to be saying that the percentages will be affected by whether any of the 10 out of 100,000 who really do have the disease got a negative test. But it doesn't. 1% false-positive means that, on average, 10 people who don't have the disease will get two false positive tests. It doesn't matter what the people who really have it get.
Positive predictive value is the likelihood that, if you have gotten a positive test result, you actually have the disease. It’s calculated as TP/TP+FP.
TP, ie true positives, is a function of false negatives.
Here's a common sense explanation. Assume the false negative rate is 100% and so you have zero true positives. Then every positive test result will be a false positive, and therefore a positive test result shows that you don't have the disease.
On the other hand, assume the false negative rate is 0. Then a positive test result will mean that you probably do have the disease as long as the false positive rate isn't too high.
Remember your post I was replying to was all about the probability that you actually have the disease conditional on a positive test result. That's the question I am addressing. For this question, you must consider the false negative rate.
Specifically, why would the 10,001st test carry a 50% chance of being a false positive?
Thus, the population for tests 1..10,000 is different than the population for tests 10,001..10,100.
Edit: I misunderstood the scenario.
If we test again, and the test is "fair", then of these 101 people, we will retest 99% (or 99) true negative, 1% (or 1) false positive, and 1 true positive. So two positives, one false positive and one true positive. 1 out of 2 is 50%.
(Tho, there is a chance also that the true positive might come back as a false negative, then 1 out of 1 or 100% of positive results would be false positive)
The silly thing with this demonstration is that accuracy is poorly defined. Tests are normally conditional on the actual presence of the disease or not.