Achieving 100Gbps intrusion prevention on a single server
blog.acolyer.org
blog.acolyer.org
* They're talking about 100Gbps ethernet traffic. Where does one even see that line rate? It seems those rates would only occur at carriers, optical transport, or in very large data-centers.
* Intrusion detection, AFAIK, would mean that the packets have to be parsed and inspected. But most traffic is encrypted, right? How far beyond typical firewall capability can you get with this stuff if you can only look at the unencrypted part of the packet?
* Not sure exactly what "intrusion" means in this context? What are they looking for? What are some common rules that get applied? Every protocol is going to be wildly different and working at an application level. Don't you need specific application experts to even think about writing such specialized rules?
* There are commercial products which, I think, do this stuff from Gigamon and Netscout. These products have special FPGA-based switches that have many 10G, 40G or 100G ports, how does the tech in the article differ? What's the use case?
In particular, I am curious about encrypted traffic (eg TLS/https/ssh). Wikipedia says IDS's can't work on encrypted traffic. "...Encrypted packets are not processed by most intrusion detection devices..."
But most traffic IS encrypted. So where do you ever have a situation where you have 100Gbps on a line, unencrypted? What's the use-case?
For traffic inspection purposes, you can still monitor the contents of (most) encrypted packets, but have to have your own certificate on the end user's device. This is most common on corporate networks, where a company will either install their own certificate as part of their device management process or rely on third party apps like Bluecoat[1] to do that.
Once they've added their own certificate to your device's trust store, that can be used to decrypt and re-encrypt traffic in transit, without any overt indications to end users that it's being done (although you can sometimes detect if this is being done). There are ways to guard against this, such as certificate pinning[2], but that's up to each individual service setting up rather than something the end-user can do themselves.
In both cases, enterprise networks are generally one of the bigger users of such software. There are open source IDS's you can use on your home network, which won't require 100Gbps. But 100Gbps is easily achievable on an office network, particularly as you start to scale to multiple offices (and if you include intranet traffic, and not just processing outbound/external traffic).
https://help.ui.com/hc/en-us/articles/360006893234-UniFi-USG...
> What are some common rules that get applied?
This is specifically answered under "Categories and Their Definitions":
> Compromised: This is a list of known compromised hosts, confirmed and updated daily as well.
> Scan: Things to detect reconnaissance and probing. Nessus, Nikto, portscanning, etc. Early warning stuff.
> SpamHaus: This ruleset takes a daily list of known spammers and spam networks as researched by Spamhaus.
> Web Apps: Rules for very specific web applications.
But there are still ways to inspect encrypted traffic. TLSv1.2 and lower expose the SNI in the handshake. Even on TLSv1.3 there is plenty of metadata exchanged during the TLS handshake. For example, the JA3 spec [1] is a standardized hash of selected connection parameters such as cipher suites, extensions, curves, etc. It can reveal a decent amount of information about the software initializing the connection, and when compared to a list of known good or known harmful JA3's can flag connections as malicious.
[1] https://engineering.salesforce.com/tls-fingerprinting-with-j...
The higher the peak throughout, the more margin you have for user applications.
If the peak throughout is 100Gbps then it’s likely that it will only consume 1% of resources at a more pedestrian 1Gbps throughout.
However, I miss more discussion on how to increase the number of concurrent flows, which would be the more complex part. Even with no rules, they only can process 500k concurrent flows, which is nowhere near enough for a 100Gbps network. For reference, in 10Gbps networks we're seeing between four and six million concurrent flows, depending on how much is that network exposed to the Internet.
I'd like to test this new system on the cheaper Xilinx ULtraScale+ FPGA using the provided software [3], but not sure will it even work with different FPGA set up with minimum change [4]?
Another thing is that it will be interesting to test and compare it with eBPF based IPS system bypassing the kernel without the need for FPGA? It seems that for Suricata IDS/IPS it has been proposed but no performance metrics are provided of the effort [5].
[1] https://www.snort.org [2] http://hogwash.sourceforge.net/docs/overview.html [3] https://github.com/cmu-snap/pigasus [4] https://www.xilinx.com/products/intellectual-property/cmac.h... [5] https://cdn2.hubspot.net/hubfs/6344338/Resources/Stamus_WP_I...
In my company we're developing traffic capture/analysis software at 100Gbps (which is orders of magnitude faster than IDS/IPS) and in order to achieve those speeds we need fast processors with a lot of cores and quite a lot of RAM, interact directly with the NIC buffers (we use DPDK now, we previously worked with modified drivers) avoiding the kernel entirely, and we have to limit the tasks per packet a lot to the point that flow state management is pretty difficult, and TCP reassembly looks impossible. I don't see a software IDS/IPS system getting anywhere close to the performance of an FPGA.
So, Snort 3.0 uses Intel’s Hyperscan library, as per the article.
Intel acquired this when they bought Sensory Networks, Inc. Guess what SN did before they moved to software? They made and sold hardware that essentially implemented Hyperscan in FPGAs (and SRAM, on PCI boards), for IDS & virus scanning at speed. They even had a patch to add support to Snort for it. SN eventually moved to a pure software model during the 2008 downturn (which smashed a bunch of their customers plans) and ultimately sold to Intel.
End of the day it is a real bitch to do right, and it was very hard at 10Gbps, and the problems are essentially identical now (only with more stuff over https).
I have no idea if they'd use them, but Intel also acquired a heap of patents on doing exactly this using the techniques that Hyperscan implements, and while they're probably happy with you doing it in software, they're probably more diligent in fighting off hardware adversaries.
If you're still interested in hardware, there's the NVIDIA BlueField 2 and its regex accelerator: https://www.nvidia.com/en-us/networking/products/data-proces...
100k flows at a line rate of 100Gbps, suggests the average flow is 1Mbps. I find that hard to believe - A typical home or office computer might have hundreds of TCP connections open at any given point in time, yet barely be downloading anything. Most of those connections are just sitting idle.
Perhaps this hardware only bothers with keeping the 100k most active flows in the reassembler... In which case the obvious attack is simply to send some packets out of order by a few seconds to make sure you're out of the most-active reassembly window?
If the OOO Engine detects that BRAM capacity for OOO flows exceeds 90% of its maximum capacity, it drops the flow with the longest linked list
The state you keep per flow also only needs to be the length of the longest pattern, so if your patterns are short (eg. banning rude words, <script> tags, etc.), it isn't much state at all... I would assume in most cases it's just a few bytes.
hard to believe? I can confirm - tcp sessions with average flows of >1Mbps do exist.
you are right in identifying a weakness of the system but incorrect as to its significance. A drag racer cannot turn, because it has been optimized for something else.
What would happen if you manufactured lots of packets to trigger the expensive filters?
then you would effectively DoS/DDoS the IPS. Now depending on how the system as a whole works it could be an efficient way to get through with a different attack that would normally be detected/blocked by IPS.
If they're only using the Snort standard filters, that would be one thing, but Cloudflare or similar services which might actually use this hardware would probably also have a completely custom rule set. So could you somehow detect experimentally what packets trigger the expensive rules? Perhaps some kind of fuzzing attack could do that.
Very interesting though.
You'd never want to build anything production-quality with one, but they're decent enough for hobbyists.