Programmers Need To Learn Statistics Or I Will Kill Them All
zedshaw.com
zedshaw.com
I don't care what your mean or standard deviation are. If your performance isn't symmetrically distributed about the mean (and it probably isn't), standard deviation probably doesn't mean what you expect.
I only look at performance in 99th percentile (or sometimes, 99.9th percentile). If 99% of responses are being served in 100ms on my webapp, I'm happy. I don't care if the mean is 5ms with a 20ms standard deviation. As long as the performance at the 99th (or 99.9th) percentile meets a level of 'acceptability' that I have predetermined, I'm satisfied. And you should be too.
This is much easier to deal with than trying to do real math. Just make sure that no more than 1% (or 0.1%, as the case may be) of requests take longer than X.
Pre-emptive answer: I usually get someone asking me why I don't demand acceptable performance at the 99.99999th percentile or at the slowest request. It's simply an issue of resources. To get that kind of performance, you can't use garbage collection, context switches become an issue, disk buffering even comes into play. You have to basically write a hard real-time (or pseudo real-time) app, and that level of effort isn't worth it for a consumer facing webapp. If I was writing code for a pacemaker or something (and please, dear god, no one ever let me do that) it damn sure better be hard real-time though.
On the other hand, I guess this argues against smanek's points being relevant to Zed's point, since statistics aren't really involved for the same reason. You don't need stats to analyze your performance if you have all the performance data; just read it off.
(Edit: I suppose his point that you shouldn't be using averages but just thresholds is still a statistical argument, against using the wrong tool like "average" on the grounds that you don't actually want the "average".)
Facebook - one random user-page view == 63 requests (reported by firebug). That means 16 people generate >1000 requests. Even with 99.9% requests within your "acceptable range", you get every 16'th person either missing an element on a page due to some DB timeout (for example), or seeing an incomplete page for a not "acceptable" time (with 1/63'th probability of it being the main document, if we don't know which element fails).
Is that really an acceptable behaviour?
I don't believe that a requirement saying "the complete website loads in less than 1s" is affected by GC or context switches at all, tbh. If you actually tested the website with load of 10x the average expected one (over a couple of hours of random scripted walk with a natural write/read balance - testing you'd do on a service you really care about) and you still get random glitches with the expected load, then there's most likely something wrong on a higher level (disk seeks failing? shared hosting overloaded? network? etc. etc.)
Going to 99.95% not-failing connections doesn't help you much IMHO. It means that you've gone from every 16'th person having a problem to every 32'th person. It's not acceptable at all.
"stating a minimum accepted value is not the same as ignoring the rest" - then what is it if you don't care about mean and stddev? If you do care about other requests you care about the mean (or stddev, or median, or whatever other parameter) but just don't put it in a scientific way. If you say reasonable, it might mean - it doesn't time-out. That means it's below X seconds. That means with the other maximum load time constraints, you expect a mean time for Y requests to be less than Z. Saying "Assuming that 1% of your requests taking long enough the user might notice" you just put a constraint on that parameter.
(just to explain why I keep arguing this point - because that's what the post is about (kind-of), people say "reasonable time", but don't want to set the actual mean/stddev - which is what they do care about)
You then look at that breakout based on what the page is, or the time of day ect. But, from a performance standpoint the the only really important number is the first number, once enough of your site is fast enough having a few slow areas, or some issues at peak times is reasonable, and once it's slow enough for a user to notice it's a problem.
Your argument is mathematically correct, but from an engineering point of view, doesn't make sense.
Of course in theory a system which responds 99.9% of the time means that it's possible for 1 of 1000 requests to die horribly. But we are engineers. If I achieve 99.99% of my target, it means that it's highly likely that that one request will also behave with pretty good performance, but I can't be sure that it will be 100ms. It might be 150ms. Or 200ms. But as an engineer, I claim that I am not concerned and don't care about knowing precisely that number.
I think this is what the author of the original article misses completely: he is applying advanced math and refuses to consider that engineering tradeoffs are practical. Of course in theory they aren't. But in practice they are.
Maybe you don't like 99% and think it should be 99.99999%. Fine, doesn't matter. Do what makes you happy for your web app.
And besides, this "measurement" has nothing to do with the functional reliability of the app. 99.99% of requests completing within 100ms doesn't mean the other 0.01% of requests fail (as another poster noted); it just means they take longer. If you're worried about requests failing (as we all should be!), then that's a separate issue to deal with. So then you say something like, "99.99% of all requests must complete within 100ms, and no more than 0.0000001% of requests are allowed to fail on average." Or something like that.
I'm unsurprised.
Programmers have a remarkable belief that they are experts in every subject, and if you tell them that there is something wrong with the way they argue, they will argue with you about it.
The physicist repeatedly insisted (and explained to me with an authoritative tone) that Python 'has no VM as it is a scripting language'.
Otherwise, it's just easier to do samples and then collect meta-statistics (which are always normal). In that case you'll need to know error rates which require things like variance and std. deviation.
Another technique is to go with statistical process control theory which mostly rejects confidence intervals and instead focuses on live sampling of processes to watch for outliers that need justification. With those you're looking at std. deviation as a measurement of the range of what's been happening vs. what's currently happening.
I think for one way flow through systems this has a tendency to be true (and so is likely to arise on most test harnesses), but I recall that multi-way routing systems with feedback effects (like, say, IP networks or highways) tend to have substantially non-gaussian congestion statistics: typically worse that gaussian, actually.
Consider B. Huberman et. al.
www.hpl.hp.com/research/scl/papers/InternetCongestion/InternetCongestion.pdf
One problem with relying upon standard deviations when you are trying to design a reliable system is that the majority of the variance in many distributions is concentrated in rare events. This is particularly true with long-tails (e.g. a Pareto distribution with alpha <= 2). If you happen to have a system conforming nearly to such statistics, and you keep your eyes on the mean and the standard deviation alone, you'll end up severely underestimating the standard deviation (which might not even converge) and you could be designing a faulty system.
Of course, the real problem, how can one gain confidence in one's statistical analysis. Not so easily, really. So if there's two things people who know about statistics, it's that statistics is critical, and statistics is hard.
Take our example of looking at response time for loading a web page. There is some finite point (say, 10 sec) beyond which we no longer care how much longer it takes. So instead of considering the distribution of response times t, we consider the distribution of min(t, 10 sec). This distribution only has support over a finite interval, so its meta-statistics normalize rapidly as you increase the number of trials.
Using this will under-report the actual standard deviation in the response time (which might, as you say, not even converge), since we've eliminated extremely low probability events with very high response time, but as a practical matter this is largely irrelevant--if these events are high enough probability for us to care we'll notice them anyway. The point of this exercise is not to perfectly ascertain the underlying distribution of t, it is to develop useful predictions for system behavior in practice.
The point is not that arbitrary statistics will necessarily always be perfectly behaved (or even well behaved) on sampling data--it's that to make reasonably accurate predictions of system behavior, under certain practical conditions, these statistics are well-behaved, and an inexperienced statistician (as most people are) is less likely to make a gross error.
Your method is a safe bet as long as you can collect "enough" data to be sure that you're covering all cases you claim to be covering. That being said, power analysis and data interpretation (ramp up, confounders, etc.) are all more difficult and important than this rant conveys. One R function is not your cure-all.
EDIT: that wasn't the original, but it was the longest. This is the original (I think): http://news.ycombinator.com/item?id=48006
Does anyone have recommendations for statistics books that are actually written in English?
* Statistics; by Freedman, Pisani, Purves, and Adhikari. Norton publishers.
* Introductory Statistics with R; by Dalgaard. Springer publishers.
* Statistical Computing: An Introduction to Data Analysis using S-Plus; by Crawley. Wiley publishers.
* Statistical Process Control; by Grant, Leavenworth. McGraw-Hill publishers.
* Statistical Methods for the Social Sciences; by Agresti, Finlay. Prentice-Hall publishers.
* Methods of Social Research; by Baily. Free Press publishers.
* Modern Applied Statistics with S-PLUS; by Venables, Ripley. Springer publishers.I agree that in practice, data munging is 90% of the work and can be pretty tedious and discouraging when working with real world datasets, but if you don't know about issues like the ones I mentioned above, how do you even know where to start?
For examples of how statistics can be used to solve interesting problems in the social sciences, just browse through Freakonomics.
And yes, I understand you cannot get published in the social sciences without using advanced statistics. All this means is that the economics academia is following into the trap of Schoclastism, and insular world that uses complicated and absurd methods of research, increasingly divorced from reality.
Social science needs to get over its science envy. As I've written here before, the proper tools to use as a student of policy and sociology are the tools of product management: http://news.ycombinator.com/item?id=836196
I have not read Freakonomics the book, but I read their blog occasionally. It's interesting when they find a simple statistic that raises an interesting point.
It might be easier to explain the problems with reference to a specific example. If there is some social science paper that uses the techniques you mentioned above, and you think is particularly good, send me a link, and I can explain why I think it's likely garbage.
What I am to do is mostly descriptive statistics, but with a few simple tests (e.g. a test of variance with covariance). I've got three populations of civil war data and would like to compare features of those populations against each other (e.g. instance of civil war in period one vs. period two vs. period three). Controlling for # of states, or # of new states (say, within five years of their founding).
Most of the complicated stuff really can't be applied to stuff like data on civil wars, since the data doesn't meet basic assumptions of the models, but I would like to be able to say something more than just 'the data from period one looks different than the data from the other two periods.'
I think this is precisely why we have more advanced statistical techniques. There are ways to correct for or at least detect serial correlation, homoskedasticity, etc, all of which are probably in your data. All real world data is fucked up in some way, having a large toolbox of statistical techniques helps you cope with this fact. Maybe not perfectly, but at least you'll know where/why your model is wrong.
That's being a little unfair on the scholastics. They did a lot of important work (particularly on logic), and the renaissance didn't just spring into existence out of nothing.
Introductory Econometrics by Woolridge is a pretty readable econometrics text you should also check out, since it'll be especially applicable for social science: http://www.amazon.com/Introductory-Econometrics-Approach-App...
It is an academic book, and is more of a history book that explains the math, rather than a math text book; which is why I found it note worthy. It's also worth reading solely for the detailed citations, from which you could completely learn the core of the field by "studying the master's" original papers.
http://www.amazon.com/Lady-Tasting-Tea-Statistics-Revolution...
And for the HN folks that like to Lisp it up:
What I got out of it was: "Some people are dumber than me." This is a well-known problem both inside and outside the programming world, and unfortunately killing the people dumber than you is not the solution to the problem.
I personally take the view that you catch more flies with honey than vinegar, but I suggest the article as a whole including the title provides a general good.
His call can be generalized: "This article is my call for all programmers to finally learn enough about [insert here] to at least know they don’t know shit."
The power that we programmers have as little gods inside systems of our own making, tends to have us over-estimate how much we know. Thus our penchant for posting stuff that makes experts in various technical and scientific fields roll their eyes.
http://www.amazon.com/How-Lie-Statistics-Darrell-Huff/dp/039...
http://www.amazon.com/Cartoon-Guide-Statistics-Larry-Gonick/...
I haven't had a chance to check out the manga guide to statistics but that might be a decent introduction, as well.
http://s3.amazonaws.com/four.livejournal/20090911/benchmark....
http://developers.slashdot.org/story/10/01/09/2154224/Why-Pr...
Q = "Zed Will Kill All Programmers"
If P || Q must hold, go with P!
On the other hand, your post sounds angry, gives no insight as to why / how he lacks knowledge (which would have been interresting) and use condescending expressions like "pseudo intellectual garbage"
what could you possibly use from this article? if you're really interested in learning stats concepts and tools you can be better served by blogs with actual content.
here's a good one: http://incanter-blog.org/
Contrary to what many short-sighted "pure" mathematicians think, Statistics is hard. You can't explain statistics in a bunch of blog posts, not even if you are a true expert. You can't condense all that knowledge in a blog post, one must read the books, though painful that may seem.
No, that's a silly approach. If your, say, launching a web app advice like this [sic] is useful. You dont need or want a deep understanding, you just need the pointers to analyse stuff right.
(as it stands I never mind Zed's tone and when I first read this article it provided some useful info to me in a simple format :)
Also r.e. your first post - from what I recall Zed can be considered some sort of an expert on this stuff.
Get this: you need deep understanding of the fundamentals, otherwise you're not doing Statistics, you're doing Voodoo Magic. If you learned some Statistics from Zed's post, then you must know even less than he does. But then, one of the many symptoms of ignorance is the false belief that one knows. Do you even realize how many years one needs to become an expert on any field, especially a tough field like Statistics?
It seems your definition of "expert" is a lot higher than mine; which is fine. Lets use a different word: experienced.
However with that said what Zed wrote in that post stood up to extra research by myself and I actually understood what he had to say.
If you wish to claim it is wrong and he has no clue, fine; but can you please - for our benefit - explain why and specifically where he is wrong?
For me an "expert" is someone who has at least 10 years of experience in one field. An expert on Statistics must have carried out extensive data-analysis work, or must have published original papers. Reading books gives one some understanding, but only when one must solve a problem, does one realize that one's knowledge is utterly superficial.
Zed is no superman. His area of expertise is not Statistics, it's something else. I am not saying he's stupid. All I am saying is that he has not invested the many years of effort into studying Statistics to earn the title of "expert". By calling him an expert, one insults all the true experts out there, the one who do not write blog posts using juvenile language because, you know, they have better things to do... like spending time with their friends and family...
I do think 2 things are happening here though. I think firstly Zed's style doesn't really appeal to you - which is fine, I can see how he wouldn't gel with lots of people. As a result the second thing happens, his content seems irrelevant.
As someone who quite enjoys his approach :) I found quite a bit of useful content. I dont think the point is really to infuse a deep understanding of statistics - it's to correct some common mistakes hackers like us make when playing with stats :)
For example I scanned the blogs of those 2 names you mention and their stuff certainly seems interesting; but frankly I didnt notice anything massively relevant to things I might need to use day to day in my "startup". Zed's stuff I've already taken to heart and corrected some of my "work practices" when dealing with stats.
As to the rest I think it is just a case of differing definitions of expert :) as I said "experienced" is a better word to use - assuming our definitions match :P.
The good news is that there seems to be an opportunity here: write Statistics for the non-statistician. Not everyone who needs Statistics can take 2 years off to learn it. Back in the 1970s, Digital Signal Processing was an advanced grad course... now it's a basic undergrad course. Maybe Statistics will go the same way, becoming more and more prevalent, and less and less ivory tower.
Sometimes it's just fun to read someone go nuts about something like statistics, and it may give you motivation to go learn more about it so that you can understand the whole thing.
But Zed's article, juvenile as it was in tone, had undeniable passion about statistics. And it was enough to keep me reading through the whole article and pay attention to the examples. As it is, I will be doing further reading on statistics and so Zed's article has been incredibly useful to me.
"Maybe it's just me, but Statistics is a beautiful field, and such intellectual beauty should be all the "psyching up" one needed."
It's just you. Nine out of ten people think statistics are nothing more than boring numbers made up on the spot.
If you can't see the beauty in it, you've probably been taught pseudo-Statistics. Don't feel bad. All the Math that engineering students learn is kind of pseudo-Math. All one learns in high-school is BS. If you want to learn something, here's my advice:
i) Don't use modern, over-designed textbooks.. you know... the thick expensive ones with many colors and boxes highlighting the formulas one must memorize.
ii) Instead read the classic books from the 1960s, many of which are published by Dover. They look boring at 1st sight, but their content is rich. Another option is to get the old Soviet books from the 1960s, which tend to be forgotten gems. The Russians are the best at marrying theory with application.
Also, just FYI: I studied statistics in business school both BS and MS, so most of my background is in applied simplified stats. This includes several graduate courses in stats in the sociology department. In addition I have extensive experience since about 2000 (so 10 years) applying statistics to problems using the R language.
Now, when you write this blog post, I'll expect to see your full Ph.D. level CV laying out all of your experience, publications, and applications of statistics just to be fair.
1) I find your writing style juvenile and distasteful. Maybe some of you yanks actually like that kind of writing. Personally, such writing makes the reading intolerable, even if the content is interesting.
2) Still regarding 1): if you already have the fame, why do you keep writing like a disgruntled teenager desperate for attention?
3) Interestingly, I don't disagree with the points you made in the article, I just don't think that insulting your audience is the best way to make your point. Maybe some people like it. To me, it just sounds vulgar.
4) Statistics taught in undergrad and business school is pseudo-Statistics. I don't mean to be pedantic, but the truth is that only graduate-level, rigorous, Mathematical Statistics counts as "true Statistics". One does not need to know the entire field (that takes at least one lifetime), but one needs to know the fundamentals very well in order to apply them.
5) Mere mortals can only be experts in one or two things. You have programming and guitar. Sorry, but you still are no Statistics expert on my book. 10 years of R is impressive and valuable, but if you don't know the foundations, you're a technician, not a statistician.
6) My field is not Statistics. I am no expert in Statistics, and I am not qualified to write about it. I can write a blog post on Algebraic Geometry, but you would not understand it, so there's no point.
7) Unlike you, I have no interest in revealing my true identity. Unfortunately, there are evil people in this world, and whatever I would write on HN under my real name could be used against me later on. If you're not in that position and can write what the fuck you want on your blog, then I must say that I envy your freedom.
To recap: writing like a teenager does alienate some readers, and calling a non-expert an expert is something to be avoided.
In this and a few other cases, the angry tone is a lightning rod that got more attention than a quiet "please guys, will you learn some stats?" If you have a better way to do it, then do it.