Applying machine learning to Infosec
conf.startup.ml
conf.startup.ml
The main differences with straight 'supervised' machine learning are the lack of labels, and the unreal volume of data (cheaply produced by machines, in machine time, we're talking microseconds to milliseconds here). So unsupervised learning is king in this domain, and often security operators have to keep an eye on and interpret the results. Another difference with other fields is that datasets are rare, mostly because of privacy as the logs can be very revealing. For this reason the market exhibits a lock that cannot be overcome by everyone I believe. Basically, you sort of need to be in the place already.
Here is a recent very high level survey on the topic (a TC report, not that bad for once), https://techcrunch.com/2016/07/01/exploiting-machine-learnin...
Here are two of our own Open Source tooling we use the most in application:
- very efficient C++ map-reduce feature generation for logs (for ML and analytics): https://github.com/soprasteria/cybersecurity-miw
- machine learning / deep learning server: https://github.com/beniz/deepdetect
The ML cybersecurity + infosec field is still young, but moving very fast, a lot of new startups and (somewhat opaque) products.
There are some reports from Google as well that you may find around, and at least at the time (few years back) they make the same observation we do with my coworkers: the rate of false positive remains high, and thus it is the job of trained security operators to spot the sophisticated attacks.
Empowering the operators with more data mining tools is a good step not yet fulfilled I believe.
Unrelated, but that gave me a thought to ponder: Will it ever be possible to use computers made inside this universe to model a MORE complex universe?
Has that already happened?
A Turing machine can simulate a python interpreter.
Have you seen the xkcd with the rocks in the desert? https://xkcd.com/505/
So, therefore, the smallest possible complete description of the state of the simulation is not larger than a complete description of the machine doing the simulating.
So if that is what you mean by more complex / what you mean by more information, then, no, the less-information thing cannot completely simulate the more-information thing.
An inner universe's state could have an initial input of some string describing all of human history up to T in the outer universe, the same differential equations, and the same amount of time T to simulate.
One might say that the inner universe contains more information, because the smallest unambiguous description of it is larger than the smallest description of its parent universe. However, the outer universe is able to simulate the inner universe.
You could mark a time and an algorithm for accumulating information, but you now have 2 more pieces of information in the inner machine than the outer one. So the outer machine has somehow managed to simulate the inner machine even though it has less information.
Although... the moment you actually carry out this experiment, it falls apart: Instead of using T, you can just look at time T + U, where U is the amount of time needed to simulate the inner machine. This is because we're sitting here talking about the experiment and eventually carrying it out, so those initial conditions encode the conclusion of the experiment.
This is subtle enough I'd have to look at the math. I can't think through it.
If yes, then can we add something to that simulation, like an extra type of particle?
- There are no publicly available data sets for training available. There are a few small ones and a few old ones, but they don't reflect the reality of 2016. Companies that approach me and pitch me solutions to the malware of 2012 are not useful.
- The majority of mobile malware is based on some kind of social engineering. On a code level these are indistinguishable from legitimate applications (the same APIs are used in the same fashion). The only difference is whether app behavior meets user expectations or not. Making this decision automatically seems intractable so far.
- Malware is not really a well-defined term. There is phishing, toll fraud, Trojans, privilege escalation exploits, ... If you generically look for malware, the signals you will look for are going to approach the complete set of APIs made available by your OS. Your results will just be a giant blob where everything is connected. Pick a single malware category and focus on just that at a time. ML signals for priv esc will look very different from those for phishing.
- ML is sexy. Malware analysis is not. Startups seem to hire too many ML people and not enough malware analysis people. I've had startups pitch to me that had literally zero people on staff who knew what mobile malware actually looked like. They just did anomaly detection and then tossed the results over to my team to verify the results. That's not how it works. We're not your QA team. :)
See the Microsoft / Kaggle challenge on classifying malware families, winning solution is > 99% accuracy IIRC.
Interesting name. Reminds me of a security scheme, Symbiotes, I briefly evaluated on Schneier's blog. Injected security into legacy, embedded applications with various tradeoffs. Where did you get the name from?
(We do some cool visual analytics work here, including unsupervised learning / classification, and target more of the problem of "given an incident you're already investigating, what else should you now look at from across all your tools?")
ML is amazingly powerful, but if you don't have sufficient domain knowledge, or you aren't collaborating very closely with actual experts, you can make very dangerous mistakes. Domain knowledge helps a lot - not just in malware, but in biology, image analysis, etc..
I think false positives are a huge problem in the info sec context and less so in other domains. To give a made up example let's say you have ML-based intrusion detection that is hotwired to some sysadmin pager. Even if your false positive rate is fairly low (at acceptable levels in other domains) with the sheer amount of data a bunch of sysadmin calls will be triggered in sum. That is a huge human factor risk as it most likely will result in teaching admins to ignore certain alerts....which of course opens up the interesting attack vector of understanding/"reversing" the ML and triggering false positives a bunch of times until you want to attack through a channel that will trigger a similar response.
Regarding the available data sets: I suppose honeypots/-nets could be created for the purpose of gathering some data sets.
Training a machine implies some sort of evolutionary model (a training set describes a fitness landscape). Maybe this will work (doubtful, across such a large and variable surface), but how about thinking about this on a fundamental design level? How does a computer know it is working properly?
That is the crux of the issue. Any command that potential malware may give, may also be given legitimately. How one tell those two apart is context. And context is a hard subject even for humans.
Even biology can't get it straight. After all, some of our most resilient diseases exploit the normal signals of cells for their own purposes.
Remember it's not infallible and should not be treated as such.