The RETE algorithm was used in a project for some work in artificial intelligence expert systems for server farm and network monitoring and management at the IBM Watson lab in Yorktown Heights.
The work started with the programming language OPS5.
Then the project invented a new programming -- language Yorktown Expert System Language 1, YES/L1. This language was a pre-processor to IBM's programming language PL/I. The pre-processing made use of the DeRemer LALR parsing tools. So, YES/LI was compiled.
And YES/LI had extensions for artificial intelligence expert systems based on the C. Forgy RETE algorithm and its working memory.
IBM shipped YES/L1 as the IBM Program Product KnowledgeTool.
Applying KnowledgeTool, we did joint work with GM.
We published several papers and gave one at a Stanford conference of the AAAI IAAI on Innovative Applications of Artificial Intelligence.
The work of the project continued with Resource Object Data Manager (RODM). RODM was more infrastructure for monitoring and management of server farms and networks.
The architectural idea of RODM was:
managing systems <--> RODM <--> managed systems
where the <--> is for two-way communications.
So, the role of RODM in the middle was something like the role of relational database between an applications program and a computer operating system file system, that is, to provide infrastructure, functionality, and tools to make the work easier, better organized, more reliable, etc.
Or, the managing systems didn't communicate directly with the managed systems but only through RODM.
So, in particular, RODM had, similar to relational database, transactional integrity. Also RODM had monotone locking protocol with automatic deadlock detection and resolution with roll-back and restart, etc.
RODM was object-oriented with an inheritance hierarchy. Each RODM object had data fields and also methods. Some of the methods were change methods and would fire (execute) when an associated field changed.
The object hierarchy was not static but was dynamic, that is, could change during execution, that is, during real-time monitoring and management. Or, the hierarchy was dynamic something like an operating system hierarchical file system that could change during execution of the operating system.
So, to keep track of the names in the object hierarchy, there was use of
Ronald Fagin, Jurg Nievergelt, Nicholas Pippenger, H. Raymond Strong, 'Extendible hashing—a fast access method for dynamic files', "ACM Transactions on Database Systems", ISSN 0362-5915, Volume 4, Issue 3, September 1979, Pages: 315 - 344.
The associated dynamic storage management was based on Cartesian trees.
The intention was that the methods could be written in rules in KnowledgeTool.
So, broadly, at the architectural level, RODM presented to the managing systems a model of, abstractions of, the managed systems with infrastructure, tools, handling the real-time aspects, coordination, e.g., via transactional integrity, etc.
Later RODM was shipped as part of IBM's NetView systems management product.
The most common approach to system monitoring was thresholds on variables with values from the managed systems, say, CPU busy, page faults per second, disk storage used, database transactions per second, network data rate.
Then essentially necessarily, monitoring with such variables is a case of statistical hypothesis tests with rates of false positives (false alarms, Type I error) and false negatives (missed detections, Type II error).
Here what statistical hypothesis testing calls the null hypothesis is an assumption that the target system being monitored is healthy, and when the hypothesis test rejects the null hypothesis is a detection of something wrong with the target system.
Such a statistical hypothesis test has higher power, such system monitoring has higher quality, when, for a given rate of false alarms, the rate of missed detections is lower, that is, the detection rate of real problems is higher.
In principle the quality can be improved by processing the variables jointly instead of separately. For this, we need a statistical hypothesis test that is multi-dimensional. Commonly a statistical hypothesis test knows and exploits the probability distribution of the input data assuming that the system is healthy (the conditional distribution given that the system is healthy). Alas, in the context of system monitoring and management, with multi-dimensional data, in practice, knowing this probability distribution is asking too much. So, we need statistical hypothesis tests that are distribution-free, that is, that make no assumptions about probability distributions.
So, for improving the methods in the objects in RODM for system monitoring, I invented a large collection of such statistical hypothesis tests that were both multi-dimensional and distribution-free.
So, in the sense of rules, expert systems, and the RETE algorithm, might change a rule from
When CPU_busy > 95% AND paging_rate >
100 AND transaction_rate < 10 Then
system_thrashing = TRUE;
to just a
three dimensional
statistical hypothesis
test on the three
variables
CPU_busy,
paging_rate,
transaction_rate.So, in this way are doing what has been called behavioral monitoring where we look for what is unusual.
At
http://www.sans.org/resources/idfaq/behavior_based.php
is
"IDFAQ: What is behavior based Intrusion Detection?"
by Herve Debar, IBM Zurich Research Laboratory with
"Therefore, the intrusion detection system might be complete (i.e. all attacks should be caught), but its accuracy is a difficult issue (i.e. you get a lot of false alarms)."
Well, commonly with a statistical hypothesis test and with the tests I invented, we get to select the false alarm rate in advance, in small steps over a wide range, and get that rate essentially exactly in practice. So, such tests solve the problem of
"you get a lot of false alarms".
That is, can set the false alarm rate as low as one pleases and get that rate in practice.