Fwaf – Machine Learning Driven Web Application Firewall
fsecurify.com
fsecurify.com
>>> p=lambda x:lgs.predict(vectorizer.transform([x]))
>>> p("/product.php?name=etc")
array([1])
>>> p("/login.php?name=rfoo&pass=hehe")
array([1])
>>> p("/download.php?file=/root/.bashrc")
array([0])
>>> p("/example/test/q=" + lorem + "<script>alert(1)</script>") # len(lorem) = 4488
array([0])
(FYI 1 means malicious and 0 means clean)https://web.archive.org/web/20170514081124/http://fsecurify....
Apologies for inconvenience.
Thanks
Another thing to keep in mind is that an accuracy of 99% doesn't mean much in an unbalanced problem like yours (much more clean queries than malicious once).
What you should show instead is precision (of the ones labeled malicious, how many are actually malicious?) and recall (out of the malicious queries in the dataset, how many did your model label as malicious?)
Going above 3 adds very,very little precision while guzzling space. Using 2 shows a drop in precision.
Probably has todo with how human built systems are built (we make them in our image - and we seem to have a thing for 3s - map(subject, verb, object) -> (origin, data, destinstion) etc)
Addendum; if you know what PCA is, id wager that added n gram dimensions share a linear dependency with lower dimensions - so sharing a statistical resemblence (covar(A,B) -> 0) that adds very little to the data's variability once you start adding dims above 3.