Corundum: Open-source, high performance, FPGA-based NIC
github.com
github.com
Corundum has several unique architectural features. First, transmit, receive, completion, and event queue states are stored efficiently in block RAM or ultra RAM, enabling support for thousands of individually-controllable queues. These queues are associated with interfaces, and each interface can have multiple ports, each with its own independent scheduler. This enables extremely fine-grained control over packet transmission. Coupled with PTP time synchronization, this enables high precision TDMA.
What makes Corundum special? Mainly that you can add client code into it and maybe free up an isolated core, or stage packets to send on short notice.
Mainly of interest for low-latency finance and for high-performance compute clusters.
Emulating such services in a packet-switched network requires very precise timing hence the reference to it in terms of IEEE1588 / Precision Time Protocol.
Supporting critical traffic with TSN is a two steps process. First, you synchronize all the participating network nodes. For this you can use PTP (IEEE 1588), which is like an Ethernet level NTP (grossly oversimplified, but you get the idea). Once all the nodes are in sync, they can use time aware scheduling (TAS) where a TDM frame is overlaid over all the LAN, with Ethernet traffic classes (TC) assigned to specific ranges. In other words, you define a repeating pattern, split into different sequential zones, and TC are aligned to some zones. The goal is to define repeating ranges dedicated to specific traffic classes, where one can control the load and make sure there is no contention and traffic will go through with deterministic latency.
All this could be used in a plant, to support both best effort traffic but also sensitive real-time traffic for automation, while protecting the later.
TSN started out for media applications (broadcasting) over Ethernet, but is getting into industrial applications (see https://opcfoundation.org/).
Support for TSN is planned for 5G (NR) release 16, to support industrial applications.
All this area is in flux, so having a flexible programmable platform can be interesting.
TDMA is basically a simple demonstration scheduler that enables and disables queues on microsecond timescales, based on PTP time. One of the original reasons for building Corundum was to enable optical switching research, where data transmission into the switch must be precisely coordinated with the configuration of the switch itself. We have tried to do this in software, but the precision is limited and the CPU overhead is high. With corundum, the schedule is enforced in hardware, so it is extremely precise and does not add any CPU overhead.
"Corundum is being developed to facilitate optical networking research and as such has some unique architectural features. First, all hardware queue state is stored in block RAM or ultra RAM, enabling support for thousands of independent, hardware controllable transmit, receive, completion, and event queues. This enables fine-grained hardware control over packet emission on a per-destination or per-flow basis. Additionally, the NIC supports multiple ethernet ports per interface that have separate schedulers but share the same hardware queues, enabling functionality such as striping packets across ports or rapidly migrating flows from port to port. The port schedulers can be made aware of PTP time, enabling high-precision TDMA that's synchronized across a large network."
https://www.reddit.com/r/FPGA/comments/cs87h1/corundum_opens...
Simple matter of some verilog and a linux driver. ;-)
Of course, we also have non-FPGA Smart NICs from the likes of Netronome, etc. which can do things like accelerate EBPF or run P4.
NetFPGA does have a NIC reference design, but AFAIK it's just the Xilinx XDMA core connected to a Xilinx 10G MAC. No accessible transmit scheduler, no offloading of any kind, etc. Just about as spartan as you can get, and it's built from completely closed components so you can't really make many modifications to it.
For what we're doing, we can't use any existing commercial NICs or smart NICs because they can't provide the precision we need in terms of controlling transmit timing. We don't care about EBPF, P4, etc. We care about PTP synchronized packet transmission with microsecond precision.
The applications for these kinds of things range from SDN (software-defined networking) where low-latency is a concern and to applications in network monitoring. One could, for example, put together a system that performs line-rate TLS decryption at 10Gbps. You need an FPGA (a big one) for something like that.
There are commercial vendors for this kind stuff (selling closed source IP and hardware). It is not yet in Open Compute networking projects, but I expect that's coming soon.
You can now buy "whitebox" switches that run open network linux and put your own applications on them. In the not-too-distant future those "applications" will also extend to stuff that can run on FPGA hardware .
It's hard for me to see the use case of an FPGA nic. The reasons outlined above don't seem compelling when a commodity nic like mellanox do so much more already.
This is could be useful for people doing testing and benchmarking on network appliances.
Corundum was originally geared more towards optical circuit switching applications, but it's certainly not limited to that. Since it's open source, the transmit scheduler can be swapped out for all sorts of NIC and protocol related research.
SDN isn't the answer either unless these FPGAs can be used directly in production and so there's a path for network cards to no longer being built on dedicated hardware. So to clarify my question I could see this being:
1) A pure research effort on network card hardware design. Useful to test things in a lab and publish papers.
2) Something that can be pushed into production by actually shipping an FPGA in the router, perhaps in specialized situations where the fixed hardware isn't flexible enough.
3) A step before actual hardware can be manufactured, and network cards themselves become a whitebox style business where multiple generic vendors show up because the designs are open-source.
Either is interesting.
It is still in development; not sure if I would trust it yet for production workloads. We will not be producing hardware; the design runs on pretty much any board that has the correct interfaces, including many FPGA dev boards and commercially available FPGA-based NICs such as the Exablaze X10 and X25.
There are, however, several limitations to it. Clock cannot be changed, for example, and usually neither can I/O, specially high-speed transceivers. This has been improving (Ultrascale Xilinx allows for reconfiguring I/O), but you still have to reserve area for reconfiguring (meaning literal area, as in a geographic region in the FPGA).
However, I/O versatility as you suggested has very few advantages to it. You need the reserved logic for ethernet to be programmed at when you plugin. Why would you leve it unprogrammed? If is simply disabled, it won't use any extra power resources, and your soft-CPU's won't be able to take advantage of these resources while you are working. Maybe you could use the area for new soft-cpus, but then you'll hit the problem of over segmenting your design and allowing for less optimization. This would inevitably impact timing constraints and area usage.
Also, FPGA programming may take minutes to finish, and always at least a few seconds. This will be very noticeable by an user and not very efficient if it has to be done frequently.
There are, of courses, good uses for that. But there is also a lot of effort on doing it right and you always risk overdoing it.
Also, for some designs you can mitigate the reconfiguration time issue by having two regions and draining requests to one of them, before doing an update. Most of the Xilinx tooling for OpenCL does this kind of thing by default (4-6 "opencl kernel" regions.) But of course it's not always an option to give up that much space...
libexanic is a remarkably clean user-space kernel-bypass library that allows you to do processing on early fragments of the packet while the rest are still being received.
[0] https://www.shi.com/Products/ProductDetail.aspx?SHISystemID=...
But Netronomes are remarkably cheap, and you can "offload" eBPC to run on the card's custom cores.
https://www.cdw.com/product/exablaze-exanic-x10-network-adap...
Ultrascale. Significantly cheaper. Could it be made to work out is it subject to your comments concerning kintex, pcie straddling?
Aside from that did you know you were going to do this when you did verilog-ethernet etc?
I would argue that it strongly depends on the price differential and how resilient your system is. If your system is properly failure tolerant, and you can buy twice as much hardware for the same price by accepting a 20% failure rate (say), then it would be strongly advantageous to buy all used hardware.
There are boards with 4 sfp+, ones with 2 sfp+ and 2 QSFP+, and even one with 4 QSFP28 (and UltraScale+ XCVU9P)...
https://www.aliexpress.com/store/group/FPGA-DEV/620372_25030...
they sound like great targets for your work...
Straddling is an attempt to mitigate this issue. Instead of only staring packets in lane 0, the interface is adjusted to support starting packets in several places. Say, byte lanes 0 and 32. Or 0, 16, 32, and 48. Now, when you have a packet end in byte lane 0, you can start the next packet in the same clock cycle, but in byte lane 16 or 32. This increases the interface utilization. The trade-off is now the logic has to deal with parts of two packets in the same clock cycle, and it has to deal with multiple possible packet offsets.
The specific annoyance with PCIe packets is that the max payload size is usually 256 bytes, but every packet has a 12 or 16 byte TLP header attached, which really screws things up when combined with the small max payload size.
Right now, 40GB is the sweet spot in lower cost surplus hardware: E.g. you can get Arista DCS-7050QX-32 for about $500 shipped on ebay all day long.
100GB/25GB switches are still really expensive.
Funny you mention that switch, we bought one of those off of eBay for our testbed as it supports PTP.
Also, for optical switching applications, one of the most important factors is how long it takes to bring up the link after switching. Because of this, we have no interest in spending time on 40G and 100G interfaces because interlace deskew takes hundreds of microseconds, and 100G also requires FEC which takes hundreds of microseconds to lock. So we're focused on 10G and 25G and running multiple links in parallel, which also provides more architectural flexibility. I added 100G support for three main reasons: the CMAC license is free, so why not?; supporting 100G makes the project a whole lot more interesting than only 10G or 25G, and it provides a simple way of testing the core NIC datapath.
Got any pointers to the sort of optical switching components you're using?
[I've been out of the networking business professionally for almost a decade now, so I'm a bit out of touch with the state of the art in optical stuff--- I was somewhat surprised recently to learn of the existence and low cost of LR4 40gb optics. :P]
Take a look at: https://circuit-switching.sysnet.ucsd.edu/
And: https://arpa-e.energy.gov/sites/default/files/UCSD_Papen_ENL...
The current generation of switches that we're working on uses diffraction gratings patterned onto glass hard drive platters, installed in a modified hard drive, spun by a custom motor controller that's synchronized to the NICs via PTP.
The cost of switch ports and interconnects could all be dumped into making host interfaces faster, allowing for the switching time to be reduced.
But, unlike portable C code, to run designs like this on real hardware, you need to do things like describe how the physical pins on the FPGA are connected to the board peripherals (for instance, describing which pin might be connected to an LED, vs a UART). This generally requires a small amount of glue, and depending on how the project is structured, some amount of Verilog/VHDL code, as well. It's not like saying "cc -O2 foo.c" with your ported C compiler that has a POSIX standard library.
This is just the case if you're using the same base FPGA, but with different board layouts. Using different FPGAs (for example, a current-gen FPGA by vendor XYZ, vs XYZ gen N-1), or especially when porting between vendors -- the details can become vastly more complex very quickly.
Very annoying for geology-hobbyists.
FWIW: this appears to be different / unrelated to ruby or rust.