Nelson Rules
en.wikipedia.org
en.wikipedia.org
For a first cut:
- Rule 1: 0.3% of samples are more than 3 standard deviations of the mean.
- Rule 2: 1/2^8 = 0.4% chance the previous 8 points were on the same side of the mean as the most recent one.
- Rule 5: 2.5% chance of being above 2sd on either side, 3 choose 2 is 3, 2 sides, so 0.375% of exactly 2. "2 or 3" is not much higher.
- Rule 6: More than 0.55%, if I've done my maths right.
- Rule 7: 0.3%
I guess you're going to get a lot of false positives if you're sampling reasonably frequently -- maybe one in 50?
When do you generate an alert? I'd say a false positive once a month would be acceptable. That'd be around 6-sigma confidence if you measure every 5 seconds.
You want a high Pr(problem|alert) but also want a high Pr(alert|problem), the trade-off you choose between them depends on how often you expect problems to occur. If they are rare then you want false positives to be rare. If problems are frequent then false positives can happen more often without affecting Pr(problem|alert) so much.
Of course, nowadays we can get so much data that we can create a procedure where an event with 10^-3 chance is seen several times a day.
A couple of rule descriptions were ambiguous, rules 5 & 6 (at least to me).
I'll upload a PDF output when I get latex installed again (downloading the html file is probably the easiest way of quickly seeing the output)
Edit - updated.
PDF:
http://files.figshare.com/2196604/Analysis_of_the_false_posi...
PDF/HTML/notebook
Lots of real-world samples follow the normal distribution, and anything that does should look roughly like that sim.
So there's no real need to use random numbers, but it's a very quick way of me getting data that looks like real data and I know I've got the standard deviation & mean correct and that there should be no anomalies.
My sim can only show one side of the story though, it can't show how often real issues are picked up. For that, we'd probably want to look at real-world data and investigate each reported issue to see what proportion are important (and then possibly try and see how many were missed).
The Nelson rules are basically an attempt to determine whether relatively 'healthy' data is actually coming from some specific forcing events (e.g. oscillation instead of random variance), so it looks for breaks with normal distributions.
http://files.figshare.com/2196656/Analysis_of_the_false_posi...
Same DOI, new version.
My guess is that positives from these rules would be logged rather than immediately reacted to. If over 1000 detections you break each rule at about the expected rate, all is well. If you break a few rules well outside of expectation, it raises some questions.
(I'm taking some rigorous approach here. By intuition it seems to me that Rule 5 should generate more false-positives than other rules)
If you are already paying $60K per year to your QA technicians, it does not matter if the rule fires once per hour or once per week. You want those guys to investigate every incident. Most of the times there will be false positives but every other month QA will catch enough minor defects to pay for itself. The real benefit though, is to catch the bigger multi-million-dollar-issues that come every other year.
[0] https://en.wikipedia.org/wiki/Common_cause_and_special_cause... Common causes are addressed by fixing the system as a whole, while special causes are addressed by fixing one-offs. For example if you are addressing variations in delivery times, you address small delivery variance by improving the maps, which help every delivery. You address the one-off variation by firing the guy who takes 3 hour lunches mid-delivery every few weeks.
In any of these cases, the Nelson Rules state that the signal is unlikely to be a stochastic random variable around the given mean. It is instead likely to have some underlying shape that is not being described.
Which includes my favorite: "Once is an accident. Twice is coincidence. Three times is an enemy action."
Even the author of the rules? Hmm…
GCSE Statistics (UK school exams at 16 years old) teach a simpler system of process control rules, closer to Western Electric https://en.wikipedia.org/wiki/Western_Electric_rules: and that is the only place I have ever come across them.
Is this in current practical use?
This is one quarter of how W. Edwards Deming promoted organizational quality control—understanding how variation works, period. (The other three being understanding psychology, understanding systems, and understanding the theory of knowledge or scientific method).
This applies directly to understanding whether observed variation has a common cause (is a natural pattern of the system), or is special cause (something unexpected): https://en.wikipedia.org/wiki/Common_cause_and_special_cause... and this impacts how you handle the variation.
For those criticizing validity, I'll say this is a way to mentally model how to understand variation, and is not meant to be 100% accurate. You're trading intuitive modeling for perfect math. But it will allow you to get close in a back-of-the-napkin quick way so you can identify patterns to study in more depth. Also, think of this in the context of many types of systems, not just a tight electrical signal pattern (which are easy systems to understand). Systems of people doing software development, machines in manufacturing processes, complex network error patterns, etc etc.
People don't often have a good idea of what's important and what's noise, especially when you don't even have a control chart but are just using intuition and a few data points. We see outliers and variations all the time in processes, especially in human processes like those we encounter in most software companies. Estimation and delays, developer performance, load failures; all kinds of complex systems that exhibit variation that people are usually "winging it" to understand.
Instead of understanding the variation and the data, people often handle every large variation in the same way, trying to "fix" it or peg it on some obvious correlation they think they observe. This says: hold on, understand what you're looking at intuitively first. Then gather more data. Don't act without understanding. Deming was fond of saying, "don't just do something, stand there!" Lots to be learned from that, and much to be gained from the simple intuitive understanding of patterns in variation.
I guarantee there are better methods.
This is already an adhockish simplification of something like Mandelbrot's seven regimes of randomness [1], which is itself, well, an oversimplification of his own work. But it formalizes the insight that quality-control is trying to impart -- the identification of inconsistencies among consistent variation.
So let's run some simulations in Matlab. We'll generate M numbers distributed like Beta and N like a Pareto (a "long tail", "black swan" distribution) with identical mean and standard deviation, and shuffle them before we interpret them as a time-series. Then we'll check Nelson's rules. Since we know how many ordinaries there are, we have a target.
In 10^4 repeated simulations each with samples of 180 ordinary betas and 20 Paretos, we expect to identify 10% of abnormals. Now, my samples are shuffled, and Nelson rules rely on time-structure (but this is precisely their weak spot); my code [2] also has visible bugs I didn't bother to fix because they'd involve thinking too hard and didn't seem so serious in large samples. Still, here are identification rates:
- Rule 1: 0.5%. This will be counter-intuitive to students of the normal distribution, but recall that the ordinary observations have a bounded distribution and we're really catching only the abnormals. Now, we've missed a lot of the 10% by this rule.
- Rule 2: 6.32535%
- Rule 3: 3.4421%
- Rule 4: 4.484%
- Rule 5: 27.83735%. This is the "medium tendency for samples to be mediumly out of control".
- Rule 6: 8.49465%
- Rule 7: 4.38775%
- Rule 8: 2.1294%
(Edited after some bug fixes that, comfortingly, didn't change the results by much!)
[0] https://en.wikipedia.org/wiki/Mixture_distribution [1] https://en.wikipedia.org/wiki/Seven_states_of_randomness [2] http://lpaste.net/137664
Then they have some sort of trend.
Then they might be autocorrelated.
Then sometimes they have some seasonality.
Also different time series are correlated between them to some degree..
If you're one of the people in the thread describing this as too subjective or strict, the MILSPEC is probably more appropriate for your process.