Cory Doctorow: Our dangerous statistical ignorance
guardian.co.uk
guardian.co.uk
It's been a rousing success, too - it's allowed our Ministry of Truth to claim things like "comrades! Our production of students that pass the standardized test is up 85% this quarter!".
It's all about families, man. Good families teach people what's really important - not federal departments of anonymoustaxeatingbureaucrat.
But yeah, he's got my vote too...
No, it means that the "99% accurate" test is wrong 9,999 times out of 1,000,000. It would be clear to anyone when stated that way. What's counterintuitive is the author's statement of the result, not the result itself.
Basically, an extremely low rate of incidence (in Cory's example, 1 in a million) will render a "highly accurate" test meaningless.
That's a huge increase from a 0.0001% chance of having the disease, but it's still not flat out terrifying. Repeat testing can weed through the false positives at a speed proportional to its accuracy.
If the test is picking up something in the person being tested, then yes, you'll get the same result every time and repeated testing proves nothing. But you can still repeat using other tests.
If the test gives false positives purely at random, then repeated testing will help. Say the test is wrong 50% of the time, and you do the test five times. If you get the same results every time, then you can be 100-(50/100)^5*100 = 97% sure of the results.
Maybe that statement is true in today's world. But for tens of thousands of years, while our brains evolved, I would guess attacks from strangers were a lot more common.
For any given student applying to some schools 1 thru n, the goal is getting at least once acceptance, and applying to more schools, mathematically, can't possibly hurt in the closed case (neglecting social engineering, time spent on applications, etc), so the chance of acceptance, Ca, approaches 1 with every new application in the following fashion:
Ca = (1 - (1 - Q1)(1 - Q2)(1 - Q3)...(1 - Qn))100
There's a visual and some worked out examples here: http://nielsolson.us/MedSchool/
Similarly, if the sensitivity of a test is 90%, that means the test identifies 9 of every 10 people with the diagnosis. If I administer n different tests each with a sensitivity of S, then the chances of accurately diagnosing the disease, Cd, goes up* with each additional positive but never gets to 1.
Cd = (1 - (1 - S1)(1 - S2)(1 - S3) . . . (1 - Sn))100%
So lets say you are doing, say, genetic testing, and any one gene is 1% sensitive for the disease. If you tested 300 genes you could be no more than 95% certain of the diagnosis.
(1 - (1 - .99)^300)100 = 95.09591...%
Now, if your genetic tests were 5% accurate, you're panel could be no more than 95% accurate with 59 tests.
If your test was 50% accurate, you're panel could consist of 6 tests and be no more than 95% accurate.
Of course, if some of the tests are negative, things get more complicated. One of the problems with these data sets is that we have no idea how predictive they are. You can't even calculate the predictive power of the database. There simply haven't been enough events. Then we get into surrogate measures (how many were positive on tests 1 - n and were found to have razor blades in their homes, etc).
The claim that these databases can't be effective isn't true. They could be. P might also equal NP. Whether the hypothesis is strictly true or not, the vague but real set of 'practical concerns' suggest that the truth of the hypothesis is sufficiently difficult to test as to render the null hypothesis the de facto assumption until proven otherwise.
1 - (5^6/6^6)
Buy 12 bars, and there's still a more than 10% chance I won't have won yet...most of the chocolate-buying government-voting lottery-praying public would be stunned.
You're probably right, and I hope that you are. Still, funny to note that Mars has even put this disclaimer on the bottom of their promo site!
(Proof left to the reader as an exercise.)
Of course, that assumes a reasonable level of intelligence, education, and drive.
Startups are about the best game going, as far as I can tell--I wouldn't be playing if the game was rigged against me (more than a little, anyway...sure, small companies have higher relative regulatory burden, but on the whole the technology game is actually rigged in favor of new companies, from a growth perspective).
However, I'm going to guess that you mean every time you try is going to influence the next time you try for the better (as evidenced by some paper about higher success rates for 2nd+ time entrepreneurs I remember on here), which makes sense. You learn from your mistakes, you make contacts, you have a better view of the market--so it shouldn't stay one in four every time you try.
So, no matter how many times you roll a die, you've only got one in six chance of rolling a 1 in all of the rolls?
Somehow, I think your math is slightly off.
But it seems that you mean, what's the chance that given x number of rolls, the very last one is a "1" (assuming you only need/want to get rich once). As the number of rolls increase, the chance of that scenerio (a string of non-1s with the last one being 1) becomes smaller and smaller when taken as a whole.
But, I'm glad we're all clear now.
Seriously, though, there have been a few studies of various degrees of reliability that indicate that new technology business failure rate over five years is quite a bit smaller than the old "9 in 10 startups fail" wives tale would have us believe. I wouldn't put significant weight on any particular piece of data, but it seems to be pretty consistently in that range whenever people who I would trust to know the numbers (VCs, angels, journalists covering the field, successful and famously unsuccessful entrepreneurs) talk about it.
And, among my own peers here in the valley who started during WFP07, about 1/7 of them are already rich (by some definition of rich). There were 21 groups, and I believe 3 have had exits, and I'm certain that Octopart, Weebly, Buxfer, Heysan, and Virtualmin have not come fully to fruition yet. Tsumobi might even surprise folks, as they're still slaving away in their secret underground lab in the Balkans (Josh may have actually said, "Boston", it was hard to hear at the Startup School reception due to the size of the crowd). YC certainly makes a notable improvement in the outcomes of their startups, but it's not magical, so I don't think it's a crazy idea to look at YC startups as at least somewhat representative of tech startups in general--where "tech startup" means, to me, folks who actually file the paperwork, build something, and get it into the hands of users...until you've done that, you're just another dork with a big idea (and those probably fail at a much higher rate than 9 in 10).
Anyway, we're only a year and a half into the experiment with the WFP07 group, and I expect the numbers will probably end up in the 1/3 to 1/2 range.
What I'm saying is, they're really smart guys working in a field that I know almost nothing about, and doing work that walks a razor fine line between "research" and "product". Thus, one of their biggest problems in reaching a market, reaching investors, or reaching developers, is making what they're working on into a concrete solution to a real-world problem that everyone (or at least their customers) can understand quickly. I think they'd be a bargain for anyone that hired them (either by investing in them or acquiring Tsumobi) because they are extremely smart kids with huge ideas, but I'm not sure how many people will see that based on what they're building.
And, while I'm pontificating, I don't think I'd be crazy to suggest that the best thing they could do would be to get their current code into the hands of some customers--even just a few. Because nothing guides you to providing value like having customers. And the more they pay (or the more ownership they have, if it's an Open Source project), the more value they demand...and that's a good thing when it comes to finding a need and filling it.
Put simply: Hecl is a scripting language for mobile phones, which are currently a real pain to program for.
25% chance, 4 cycles: 68.4% chance of one or more successes. 33.3% chance, 4 cycles: 80.2% chance at least one is a winner.
33.3% chance, 10 cycles (your entire career): 98.3%. Though it would be pretty hard to keep at it at 58, after 9 failures.
Here tis: http://craphound.com/littlebrother/Cory_Doctorow_-_Little_Br...
I'm sure that's true, but it doesn't answer the real question. For most parents, the question is: given my child's particular environment, what is the greatest threat to him?
Doctorow is saying: given that your child was attacked (and no additional information), the attacker is more likely than not to be a relative. That seems backwards.
Of course, since P(Child Attacked) is so very low it's still not a huge deal.
It's not really backwards. That sort of inversion is exactly part of Bayes' Theorem.
Picking numbers from thin air:
Let's say are related to 30 people know 200 and you encounter 10,000 random strangers and you have a 1 in 100 chance of an attack. Well the odds a specific random stranger attacks your child given the above assumptions is less than 1/2,000,000. The odds a specific person they know attacks your child would be less than 1/20,000 and the odds that a given relative attacks your child would be more than 1/6,000. Now who would you focus on? (What about a predator that has a 50% chance of attacking someone well he is still under 1/6000 because there are so many other people for them to attack.)
PS: Given my child's particular environment you should focus on relatives and people you know.
many if not most citizens of the USA do not understand the basic organization and functions of their government. most of them cannot name a single sitting supreme court justice, have never read the constitution, do not know how a bill becomes a law.
and these people vote.