Liberouter Combo Cards – FPGA boards focused on network data processing
liberouter.org
liberouter.org
The idea is to move "network decision making" (load balancing; maybe firewalls and other assorted stuff) from the CPU to the FPGA, where this kind of thing can be done faster, and more reliably (e.g. you can have hard timing constraints on a FPGA).
Guess the other players in this field are pure software switches, which can be very flexible, but sometimes slow. And ASICs, which are very fast (potentially faster than FPGAs) but not as flexible.
[1] - http://arstechnica.co.uk/information-technology/2016/09/micr...
Would an FPGA really give advantages over such a setup? I can't imagine it'd beat the CPU at being a control plane—so it'd mostly have to be that it has a lower TCO as a data plane than an ASIC. That might be true, if the code in the data-plane would ever need updates (which, for ASICs, mean new hardware revs.)
In this case, each customer would have a relatively small number of the switches where volume doesn't justify an ASIC. Might be enough switches & customers to justify licensing a FPGA core with development cost spread out among many customers.
ASIC folks be like, "You gonna neef to bring a soldering iron for that update. Or buy our latest model." Haha.
I've been really stymied by how to process packets in hardware. The obvious approach seems to be to run an RTOS or even full-fledged linux if your FPGA has hard IP cores on it. But is there a better way? How much performance would one lose? I'm also a bit confused about how to communicate on PCI-Express (I'm a software guy ... so learning about DMA). I have seen soft IP for TCP/IP stacks but it seems too crazy. I'm doing this as a hobby education project btw. It has been great fun so far! Wish there were meetups on this topic.
[1] Such as https://www.amazon.com/Logic-Design-Verification-SystemVeril... [2] https://www.xilinx.com/support/documentation/ip_documentatio...
https://www.juniper.net/documentation/en_US/junos15.1/topics...
You can get Xilinx's component for their pma/pcs for 10g base-r ethernet for free from vivado and stick one of the macs from open cores on the end of it (probably this: http://opencores.org/project,xge_ll_mac - I used it for prototyping and it seems to work (before creating my own pcs/pma block and mac to cut down the latency)
Once you have that, then you would need to deal with the ethernet frames streaming through the FPGA, probably 64-bits at a time at 156MHz for 10G, so you need to pull out the fields you are interested in (like mac addresses, ip addresses, etc). You can buffer the incoming packet into a FIFO whilst waiting for the stuff you want to filter on. Once you have all your fields you can decide whether you want to pass the packet through to the tx side or not (I usually read the packet out of the FIFO either way and just hold the valid low for packets I don't want to send).
Hope this makes sense!
I looked into this a few years ago. If you can settle for just sending bytes on an ethernet cable, you can relatively easily TX/RX using UDP Multicast. The FPGA can beam bits straight to the PHY adapter and out they go.
Actually, not quite, not with FPGA. Sure, you can use an HDL to synthesize a custom microchip, in which case you indeed would be "building hardware", but you can also target an FPGA specifically, in which case you would essentially be programming a kind of a computer in a way similar to how it was done on those (early) switchboard-based computers that were configured by manually making wired connections between a fixed set of logical and arithmetical devices. You could say that one would be "building hardware" in such case as well, but since the system could be easily reconfigured at any time so that it could perform a different function, to me it looks more like programming... (I guess, the right way to look at this is as something where the distinction between hardware and software gets blurred to the point that any serious argument about the meaning of words becomes, well, meaningless.)
FPGAs are a undifferentiated sea of configurable logic gates with configurable connections between them[0]; it's not quite wiring together bare transistors but it's not that far removed either. None of the elements of a von Neumann architecture are there; if a engineer wants any of those things in their design, they would have to assemble it themselves from the available gates[1]. So, no, building a design for a FPGA has little to do with programming in the sense that most people mean it.
[0] With various special function blocks interspersed at regular intervals depending on the manufacturer and model of the FPGA
[1] or buy a pre-made soft-IP core
No disagreement here, but see, e.g., http://fpgacenter.com/fpga/fpga_prg.php.
Letting the OS touch more than a small percentage of the packets would be a gross waste of resources. If you want to write a program to process packets, use a real processor. (Putting an FPGA in a computer and then running a soft core on the FPGA is even sillier for performance.)
The main reason to use an FPGA for this is to build structures that effectively encode your firewall or other layer processing rules in tables in the hardware, so the hardware can make a decision without having to consult the processor. MPLS and VLAN routing is the obvious case, but you can usually keep a short IP routing table in there too. Ideally you'd be able to make this decision early on while recieving a packet so you can start transmitting it while the last bits are still coming in. What you want to avoid is buffering ("bufferbloat").
Excess buffering is a problem, but a little bit is fine unless you're trying to do high frequency trading. If you don't move packets to an external DRAM, you're in no danger of having bloated buffers with just the memory on the FPGA.
Or you could implement an active queue management algorithm to keep your buffers from inducing too much unnecessary latency even when they are generously sized. There are a bunch of AQMs out there to pick from, of varying complexity. And it doesn't seem like there's near enough research into hardware implementations suitable for use in switches or network co-processors.
Searched a bit, and found a terrific comparison of doing the same work three different ways. It compares serving up a Key/Value store using traditional software, then DPDK, then an FPGA based approach: http://www.hoti.org/hoti23/slides/lockwood.pdf
Skip to slide 22 if you're impatient.
Custom silicon, rather than the FPGA or ASICS they were using when they started the project in 2012.
They tout not only the freed CPU resources, but lower latencies and increased security of isolating the networking from the CPUs running hosted VMs.
http://web.archive.org/web/20160604121207/https://www.libero...