AMD proposes an FPGA subsystem user-space interface for Linux
phoronix.com
phoronix.com
This has always felt like a gaping security hole waiting to be explored.
Modern, high end FPGAs have a feature known as Raw SerDes, which in essence allows you to bypass a PCIe or Ethernet controller and use those lanes (yes, PCIe lanes) to your heart's desire ...provided you can design a working communication protocol. Difficult, but not impossible by any means.
So if you wanted to, you could design your own PCIe controller and give it whatever device ID, vendor ID, memory space, or capability space you want! Normally these things are not writable on a PCIe controller. But if you designed your own, you could write them to whatever you want and spoof device types, memory spaces, or driver bindings, and probably get yourself access to memory you shouldn't be touching. While I don't know how the linux kernel would handle these potentially out of spec conditions, it never sat right with me from a security standpoint.
Not in a system with a properly configured IOMMU unit. That stuff got some serious attention back in the old Thunderbolt 2 era, when people discovered that yes, it's PCIe under the hood and yes, having no IOMMU protection yields an attacker an instant-0wn.
> a feature known as Raw SerDes
I have never heard anyone use the term "raw serdes" for hard transceiver IP cores.
And since it's near impossible to validate FPGA firmware functionality by the kernel, rights to send bitstreams to the FPGA is essentially equivalent of root on DOM0.
That's like saying "did you know that advanced microprocessors have the capability to bypass I2C and output voltages directly on the pins?!?!?!?"
First of all, it's backwards. The physical layer comes before the protocols and is always there at the base. Second of all, the world does not exclusively run on I2C. Some people want SPI busses or to toggle transistors with GPIO. That's fine. Sure, gate it behind different permissions, but don't just rant at what you don't understand.
If you want a concrete example where serdes access is important, look up JESD204b, but in general there are loads of real-time systems or bespoke processing applications where it makes sense to dispense with the complex and temperamental packet-switched infrastructure in places where that complexity and nondeterministic behavior is likely to cause more trouble than good. There are also applications to backplane connections (if you are encapsulating PCIe, you want to run slightly faster than the PCIe so the PCIe can run at full bandwidth), even to the development of next-gen PCIe itself. It's not magic, it is not delivered by a stork, it needs to be prototyped, and that's another thing FPGAs are used for.
In any regard, a lot of threat models (including mine) consider installing hardware (especially an FPGA) as a trusted action.
Otherwise if it's configfs you're root on the system and unless it's integrated peripherals you plan to attack you probably have finer grained hardware context to imply physical access... which seems to minimize the farther reaching, generalizable concerns?
Not even close to impossible. I've recently been trying to figure out a relatively "low-tech" way of talking to modern displays that doesn't involve feeding VGA into a black-box ADC, and from what I've gathered so far, most of the serial link standards developed in the past ~25 years are basically overclocked Fibre Channel with the serial numbers filed off and most of the reliability/ordering guarantees quietly dumped in a roadside ditch.
Fascinating. Do you have a blog post about this?
DRAM is still almost always a traditional parallel bus, and it's basically alone in that. HDMI was a late holdout where the signalling was serial on differential pairs but the clock speed was variable depending on the data rate; newer versions of the standard rely instead on the link operating at one of a handful of fixed standard data rates, as done by everything else.
PCI express, SATA, usb3, HDMI (in a slightly different way), display port, etc etc. all use some form of 8b10b coding (or more efficient coding) with multi-gigabit serial transceivers.
In fact the PHY layer for USB3, PCIe 2.0 and SATA is identical - Intel designed a physical standard to encompass all three of those called PIPE. Nowadays that only exists as a virtual bus between silicon IP blocks.
FPGA vendors will also justify inertia in that current FPGA users don’t seem to be deterred by the bad tools because of the economics of their business.
Some think a lot of hobby users would try FPGA if the toolset was easier to pick up but there are not enough of those folks to keep Radio Shack or even Fry’s alive and they will be buying $5-$150 parts, not the much more powerful $10,000-$100,000 parts.
This has been the persistent argument for many years from companies who say they can't release Open Source graphics drivers.
> FPGA vendors will also justify inertia in that current FPGA users don’t seem to be deterred by the bad tools because of the economics of their business.
Want hundreds as times as many FPGA users? Make it easy for an FPGA to be used for transparent acceleration, by making it easy for Open Source libraries to build and ship FPGA bitstreams that serve as accelerators for their data handling. Imagine if compression libraries, databases, and many other kinds of libraries could transparently take advantage of an FPGA if available to process data many times faster. Then there'd be a benefit to shipping an FPGA in many servers, and many client systems as well.
Those that are willing to go out of their way to design a custom circuit or something else on an FPGA are in my opinion the type already dedicated enough or driven enough to not be deterred by crappy tools.
The work you do on FPGAs is already a filter enough so that I don't think anyone is getting passed that and then giving up because the tools suck.
About 10 years ago I was doing some FPGA development in a startup. We were using Vivado. It seemed like we spent about 30% of our time working around Vivado bugs. I come from a hardware background originally and then got into software development (EDA tools) later on. After the startup gig ended I could've gone more in the direction of FPGA development. I decided not to because the tools suck and life is too short to deal with that day in and day out. And it's not simply that the FPGA vendor tools are some of the buggiest software known to humankind, it's that the FPGA vendors don't care to make them better.
It doesn't take 100x the devs to make FPGA compelling on the desktop or the server. Just like bespoke accelerators in Apple Silicon are used behind a library, so too will the accelerators implemented via FPGAs. The program itself can be copied a billion times.
Your argument can be made for GPUs as well, the users (end users) aren't the ones writing the shaders, but GPUs are used my hundreds of millions of people.
What? How can any company claim that with the patent thing at play? Wouldn't that just be admitting they're violating patents, therefore making the closed-sourceness reason moot in the first place?
Moreover, wouldn't any sufficiently-interested patentholder just reverse-engineer the compiled binary and arrive to the supposed infringment on their own?
It's a MAD (mutually assured destruction) situation. You can rest assured that everyone knows about everyone else's rotting corpses in the storage locker... the first one to chicken out to the feds will get blasted to pieces just like everyone else.
My personal opinion is that today's patent systems can go and die in a fire for all I care, right after copyright.
Example: Modern DDR5 has something like 64GB of bandwidth per channel; assuming your design is inline on the bus running at something like 500MHz, you'd need a 128-bit bus, per channel. That clock rate might require deep pipelining, further increasing area requirements, so you can't fit as much other stuff. Otherwise, you need a wider bus and to go slower, but wider buses often scale sub-linearly in terms of area and routing congestion; a 256-bit bus will be more than twice as expensive and difficult to route as a 128 bit one due to limited routing tracks, etc. So maybe you can hit that target, but then you're too routing congested, so you can't fit as many channels as you want in. Ergo, you need bigger/more FPGAs, or serious optimization and redesign. There's no immediate win. You need to explore/napkin math the design space to find the best solution on the pareto frontier, typically. Or just buy a FPGA that's massive overkill, AKA "buy a faster PC", the typical software programmer's solution. But it really isn't plug and play or anything close to that.
It's similar to other niche things, like in-memory GPU databases. They are not held back by CUDA being proprietary. That fact does suck, but it's not really relevant in the grand scheme. They are held back by physical design dictating that parallel accelerators need loads of fast memory to feed the execution units, fast memory is super expensive and takes up a lot of space on the PCB resulting in a physical upper bound on density, and that the working set for such databases typically grows much, much faster than rate at which GPU memory performance/price drops. Past the point of no return (working set > VRAM), their advantages rapidly vanish. Their limitations are in the design, not the software.
FPGAs taught me a lot about hardware/software design. I really like them and want more people to use them. I'm really excited there are fully FOSS flows, even if they have giant limitations. But they are pretty niche and have serious physical design factors to account for; I say that as someone who contributes to, uses, and loves the open-source tools for what they are, and even was lucky enough to play with them for work.
And yes, you'd need to leave it programmed with the accelerators you actually need. You could have system policy that programs in the accelerators for libraries your software uses, with a mechanism for saying "there's not enough room in the FPGA for all the accelerators, pick the ones you want".
Among other uses, this would mean you might not need specialized hardware for video decoding or encoding for each new codec; you could put it on an FPGA, and upgrade it in the future. In theory you could put it in place of more special-purpose transistors that you can run on the FPGA instead.
If you end up doing something like this, it's generally only because your workloads are extremely atypical, e.g. Google's Video Processing Units or whatever might be good as FPGAs but only because they're such outliers. Actually they're just using ASICs because that's more economical. But my point is there isn't actually enough of this to go around in a way that trickles down to the consumer.
It's similar to the question "Why doesn't my x86 CPU have 1,000 cores like a GPU" or "Why isn't all of my memory SRAM." Because it just isn't actually useful or what anyone actually wants and the costs are vastly disproportionate to the actual utility.
On FPGAs designed for this, it is possible to "gradually reconfigure" FPGAs on context switch at high speed, while they continuously process data, in a manner similar to how CPUs gradually change what's in their cache after a context switch, and modern GPUs handle multiple applications by scheduling work units across the compute elements.
I expect those sorts of FPGA designs would become available on the market if vendors decided to develop the ecosystem of FPGAs as general purpose compute accelerators, shared among applications, similar to the role played by GPUs, TPUs and NNPUs now.
(Long shot: If anyone out there seriously wanted to hire someone to build open source or open programming, high performance FPGAs with these switching characteristics, and tooling to match, I would love to do both :)
And people do actually create marketable FPGA designs you can load into modern accelerators. You can buy Bittware devices yesterday, or Xilinx Alveo and load tons of designs into them. You can go get Amazon F1 instances and put tons of accelerators on them. You don't hear about them and they aren't popular like GPUs because the fact is that most people don't need this, and the ones who do have very particular designs that probably aren't worth over-optimizing the entire system architecture for. That's why they're 95% PCIe cards with attached output peripherials that most of the time end up in Ethernet.
Those devices are completely different to use compared to the sort of general purpose, fast-compilation, fast-switching accelerators like modern GPUs.
FPGAs and FPGA-like architectures and concommittant design software can be designed for fast compilation, adaptive timing and pipelining, and overlapped application multiplexing. But it takes significant design changes. It's a novel and underexplored area. With such architectures, schlubs like us can write software that runs on them with excellent performance for some tasks.
Unfortunately the market and the legal situation hasn't optimised for that. The closed FPGA programming information, for decades, meant others could't produce radically different commercial tools for existing FPGAs, which would generally require skipping the proprietary P&R to use novel fast-compilation and incremental reprogramming techniques. Those who explored it were always worried about legal issues, as well as damaging customer devices.
And for a long time the patents were a chilling effect on new entrants wanting to develop alterate FPGA architectures better suited to this type of programming, as long as they contained elements of traditional FPGAs as well. The patent situation is starting to shift now that early Xilinx and Altera devices are old enough, but it's a multi-decade process, unfortunately.
It's pretty obvious that an FPGA is a bad choice as an accelerator if you can get away with it. Future CXL FPGAs will be highly capable platforms, but they will be both expensive and a nightmare to develop for, negating most of the reasons why you would use them.
By the way, your complaints about LUTs taking up transistors is pretty irrelevant. Most of the transistors are being "wasted" on the routing switches and connection boxes. The space taken up by LUTs is so small that there are mask programmable gate arrays, aka FPGAs without the routing switches and connection boxes. They end up three times as dense as a regular FPGA as a result.
Assuming the patentholder had sufficient and warranted suspicions, wouldn't they initiate legal action and get the actual source/hardware design files through discovery anyway?
https://f4pga.readthedocs.io/projects/prjxray/en/latest/arch...
Maybe. There certainly is a lot of "secret sauce" energy around the bitstream formats. Primarily I think they guard the bitstream format to help ensure vendor lockin. Imagine if there were open tools that could easily target FPGAs from multiple vendors so that users could choose the most cost effective solution. The FPGA vendors don't want that.
Because the word "subpoena" doesn't exist for any of these companies?
However, I don't think that this is a real issue, as competitors and the most skilled customers already mostly know how the devices work. Also, both Xilinx (AMD) and Altera (Intel for now, but looks like the might spin it out) have so many patents that it's probably mutually assured destruction and a huge gamble if either sues the other. I think they just prefer having the tools proprietary, not just to lock the customers in (though they like that) but also to keep away from the hairy corners (avoid defects in the hardware design that could produce bad results or fry the chips).
The bitstream format is an obstacle but can be reversed. It's already been done for the Xilinx 7 series, lattice ecp5, and others.
However, that does NOT solve the actual main problem - timing models.
Timing models are huge databases hundreds of megabytes for a single FPGA that provides exact routing delays and propagation delays for groups of functional gates in the fabric. They are developed over months of painstaking analysis, debugging and tooling by the vendor.
Timing models are what let's you say "please make this IP run at 166mhz" and the fitter knows exactly how hard to work placing and connecting LUTs so that is possible.
Then, the timing analyzer will check the maximum frequency of that clock domain and ensure, using the timing models, that the specified frequency can be reached at all 4 process corners (PVT).
A typical design will have usually anywhere from 5 to 30 clock domains.
So if you have no support for the timing model, you effectively are not able to ever optimize your fitting process, and you have no idea if your FPGA will have its state machines explode when it gets a bit warmer than ambient.
Theres been some incremental improvement with low level Linux support across the past year. Good to seem but so far, actually using FPGAs is still all Vivado & closed systems. I think there's so much possibility left on the table by not supporting openfpga alternatives, not embracing yosys/openpnr/openroad/&al.
GPUs are really fast but using some kind of instructions, while on a FPGA you can practically design what "instruction" is going to run. From the limited experience I had on CUDA but the complexity of your algorithm and how much branches your code have can make your code a lot slower than running it on a CPU, no matter the amount of CUDA cores you have.
It could be incredibly cool that in some near to mid future to being able to run an application that could run some kind of code in the FPGA, like they already does with the GPU, to solve some kind of problems, like audio, image, video processing, and probably machine learning (I don't know too much about current implementations), and with that user-space interface, it could even be earlier than I though.
Where FPGAs shine best are in embedded or very obscure applications where GPUs are way too big, power hungry, or cannot support enough parallelization of your specific algo.
The logic must be very carefully implemented to allow over-writing only specific regions of the FPGA.
If you're re-flashing the entire logic array, it's pretty straightforward. If you want to leave some logic in and add an additional funcitonality, this is much more difficult.
This is really a limitaiton of the FPGA logic array implementation, not a linux shortcoming...
We badly need them for the software-defined things (SD-X) namely software-defined radio (SDR), network (SDN), Instrumentation (SDI), etc, and Linux is at the forefront of the transformation.
Be very nifty if this takes off.
FPGAs are general purpose in a broader sense than CPUs or (GP)GPUs. And, no other general purpose computing device is able to transport and compute incoming data in a pipeline in a few cycles (10, 100, whatever) from input to output. 30Gbyte per second? No problem. Under certain circumstances even guaranteed on a normal PC. A general interface such as a BIOS mapped into the memory would be interesting for many applications.
Apple has taken a good step with the Afterburner (like others before it in the music and movie sector).
After the coin hype, FPGAs once again disappeared into the niche of high-frequency trading, only to become popular in the low-price spectrum of chips for console emulators in recent years. No cloud provider has managed to bring FPGAs to the masses. Intel and AMD have bought the market leaders and have blown all efforts to achieve something in the data centers.
AMD seems to take the first step: introduction of a general interface. The directly necessary next step would be something like CUDA (apart from IP, the name once stood for "Compute Unified Device Architecture") instead of VHDL or Verilog. It doesn't even have to be open source. It has to be something where you can compile a demo program (Mandelbrot and RISC V and 68030) in five minutes and get it running (and don't always tell everyone that it's called "synthesizing"). On a PC with a typical graphics card, place & route would be much faster. But not with a tcl/tk interface in Java as with Vivado.
I guess what they're talking about is upstreaming all of this...
(I believe the Xilinx-forked kernel uses a derivative of these patches)
What is FPGA for?
FPGAs are often times for prototyping hardware designs, but also for bespoke parallelized or hardware operations that are too low of yield to justify ASIC development.
ASIC -> is hardware that may run only single program which is hardcoded into the chip.
FPGA is something between CPU and ASIC.
Right? (At eli5 level)
An example of this is video decoding/encoding, which is commonly implemented by dedicated hardware.
Admittedly, I haven't had the opportunity to play with FPGAs very often in my professional career, but the limited experience I've had with them showed me you need an entirely different mindset when programming in a hardware description language. With assembly everything is global, and similarly with FPGAs almost everything is asynchronous and parallel.
I bet this is primarily to compete with fixed function AI accelerators, so that you can adapt to new generations of the AI hype.
It doesn’t really need to be a better piece of engineering, or more power efficient. I’d say it is more guided by a marketing thing.
Nvidia provide drivers for what feels like a VERY long time.
I have the feeling that AMD gives up on providing drivers after only a few years of product life.
But these are only hunches, anyone know the facts?
I don't believe that.
My question stands - anyone have FACTS?