FPGA vendors will also justify inertia in that current FPGA users don’t seem to be deterred by the bad tools because of the economics of their business.
Some think a lot of hobby users would try FPGA if the toolset was easier to pick up but there are not enough of those folks to keep Radio Shack or even Fry’s alive and they will be buying $5-$150 parts, not the much more powerful $10,000-$100,000 parts.
This has been the persistent argument for many years from companies who say they can't release Open Source graphics drivers.
> FPGA vendors will also justify inertia in that current FPGA users don’t seem to be deterred by the bad tools because of the economics of their business.
Want hundreds as times as many FPGA users? Make it easy for an FPGA to be used for transparent acceleration, by making it easy for Open Source libraries to build and ship FPGA bitstreams that serve as accelerators for their data handling. Imagine if compression libraries, databases, and many other kinds of libraries could transparently take advantage of an FPGA if available to process data many times faster. Then there'd be a benefit to shipping an FPGA in many servers, and many client systems as well.
Those that are willing to go out of their way to design a custom circuit or something else on an FPGA are in my opinion the type already dedicated enough or driven enough to not be deterred by crappy tools.
The work you do on FPGAs is already a filter enough so that I don't think anyone is getting passed that and then giving up because the tools suck.
About 10 years ago I was doing some FPGA development in a startup. We were using Vivado. It seemed like we spent about 30% of our time working around Vivado bugs. I come from a hardware background originally and then got into software development (EDA tools) later on. After the startup gig ended I could've gone more in the direction of FPGA development. I decided not to because the tools suck and life is too short to deal with that day in and day out. And it's not simply that the FPGA vendor tools are some of the buggiest software known to humankind, it's that the FPGA vendors don't care to make them better.
It doesn't take 100x the devs to make FPGA compelling on the desktop or the server. Just like bespoke accelerators in Apple Silicon are used behind a library, so too will the accelerators implemented via FPGAs. The program itself can be copied a billion times.
Your argument can be made for GPUs as well, the users (end users) aren't the ones writing the shaders, but GPUs are used my hundreds of millions of people.
What? How can any company claim that with the patent thing at play? Wouldn't that just be admitting they're violating patents, therefore making the closed-sourceness reason moot in the first place?
Moreover, wouldn't any sufficiently-interested patentholder just reverse-engineer the compiled binary and arrive to the supposed infringment on their own?
It's a MAD (mutually assured destruction) situation. You can rest assured that everyone knows about everyone else's rotting corpses in the storage locker... the first one to chicken out to the feds will get blasted to pieces just like everyone else.
My personal opinion is that today's patent systems can go and die in a fire for all I care, right after copyright.
Example: Modern DDR5 has something like 64GB of bandwidth per channel; assuming your design is inline on the bus running at something like 500MHz, you'd need a 128-bit bus, per channel. That clock rate might require deep pipelining, further increasing area requirements, so you can't fit as much other stuff. Otherwise, you need a wider bus and to go slower, but wider buses often scale sub-linearly in terms of area and routing congestion; a 256-bit bus will be more than twice as expensive and difficult to route as a 128 bit one due to limited routing tracks, etc. So maybe you can hit that target, but then you're too routing congested, so you can't fit as many channels as you want in. Ergo, you need bigger/more FPGAs, or serious optimization and redesign. There's no immediate win. You need to explore/napkin math the design space to find the best solution on the pareto frontier, typically. Or just buy a FPGA that's massive overkill, AKA "buy a faster PC", the typical software programmer's solution. But it really isn't plug and play or anything close to that.
It's similar to other niche things, like in-memory GPU databases. They are not held back by CUDA being proprietary. That fact does suck, but it's not really relevant in the grand scheme. They are held back by physical design dictating that parallel accelerators need loads of fast memory to feed the execution units, fast memory is super expensive and takes up a lot of space on the PCB resulting in a physical upper bound on density, and that the working set for such databases typically grows much, much faster than rate at which GPU memory performance/price drops. Past the point of no return (working set > VRAM), their advantages rapidly vanish. Their limitations are in the design, not the software.
FPGAs taught me a lot about hardware/software design. I really like them and want more people to use them. I'm really excited there are fully FOSS flows, even if they have giant limitations. But they are pretty niche and have serious physical design factors to account for; I say that as someone who contributes to, uses, and loves the open-source tools for what they are, and even was lucky enough to play with them for work.
And yes, you'd need to leave it programmed with the accelerators you actually need. You could have system policy that programs in the accelerators for libraries your software uses, with a mechanism for saying "there's not enough room in the FPGA for all the accelerators, pick the ones you want".
Among other uses, this would mean you might not need specialized hardware for video decoding or encoding for each new codec; you could put it on an FPGA, and upgrade it in the future. In theory you could put it in place of more special-purpose transistors that you can run on the FPGA instead.
If you end up doing something like this, it's generally only because your workloads are extremely atypical, e.g. Google's Video Processing Units or whatever might be good as FPGAs but only because they're such outliers. Actually they're just using ASICs because that's more economical. But my point is there isn't actually enough of this to go around in a way that trickles down to the consumer.
It's similar to the question "Why doesn't my x86 CPU have 1,000 cores like a GPU" or "Why isn't all of my memory SRAM." Because it just isn't actually useful or what anyone actually wants and the costs are vastly disproportionate to the actual utility.
On FPGAs designed for this, it is possible to "gradually reconfigure" FPGAs on context switch at high speed, while they continuously process data, in a manner similar to how CPUs gradually change what's in their cache after a context switch, and modern GPUs handle multiple applications by scheduling work units across the compute elements.
I expect those sorts of FPGA designs would become available on the market if vendors decided to develop the ecosystem of FPGAs as general purpose compute accelerators, shared among applications, similar to the role played by GPUs, TPUs and NNPUs now.
(Long shot: If anyone out there seriously wanted to hire someone to build open source or open programming, high performance FPGAs with these switching characteristics, and tooling to match, I would love to do both :)
And people do actually create marketable FPGA designs you can load into modern accelerators. You can buy Bittware devices yesterday, or Xilinx Alveo and load tons of designs into them. You can go get Amazon F1 instances and put tons of accelerators on them. You don't hear about them and they aren't popular like GPUs because the fact is that most people don't need this, and the ones who do have very particular designs that probably aren't worth over-optimizing the entire system architecture for. That's why they're 95% PCIe cards with attached output peripherials that most of the time end up in Ethernet.
Those devices are completely different to use compared to the sort of general purpose, fast-compilation, fast-switching accelerators like modern GPUs.
FPGAs and FPGA-like architectures and concommittant design software can be designed for fast compilation, adaptive timing and pipelining, and overlapped application multiplexing. But it takes significant design changes. It's a novel and underexplored area. With such architectures, schlubs like us can write software that runs on them with excellent performance for some tasks.
Unfortunately the market and the legal situation hasn't optimised for that. The closed FPGA programming information, for decades, meant others could't produce radically different commercial tools for existing FPGAs, which would generally require skipping the proprietary P&R to use novel fast-compilation and incremental reprogramming techniques. Those who explored it were always worried about legal issues, as well as damaging customer devices.
And for a long time the patents were a chilling effect on new entrants wanting to develop alterate FPGA architectures better suited to this type of programming, as long as they contained elements of traditional FPGAs as well. The patent situation is starting to shift now that early Xilinx and Altera devices are old enough, but it's a multi-decade process, unfortunately.
It's pretty obvious that an FPGA is a bad choice as an accelerator if you can get away with it. Future CXL FPGAs will be highly capable platforms, but they will be both expensive and a nightmare to develop for, negating most of the reasons why you would use them.
By the way, your complaints about LUTs taking up transistors is pretty irrelevant. Most of the transistors are being "wasted" on the routing switches and connection boxes. The space taken up by LUTs is so small that there are mask programmable gate arrays, aka FPGAs without the routing switches and connection boxes. They end up three times as dense as a regular FPGA as a result.
Assuming the patentholder had sufficient and warranted suspicions, wouldn't they initiate legal action and get the actual source/hardware design files through discovery anyway?
https://f4pga.readthedocs.io/projects/prjxray/en/latest/arch...
Maybe. There certainly is a lot of "secret sauce" energy around the bitstream formats. Primarily I think they guard the bitstream format to help ensure vendor lockin. Imagine if there were open tools that could easily target FPGAs from multiple vendors so that users could choose the most cost effective solution. The FPGA vendors don't want that.
Because the word "subpoena" doesn't exist for any of these companies?
However, I don't think that this is a real issue, as competitors and the most skilled customers already mostly know how the devices work. Also, both Xilinx (AMD) and Altera (Intel for now, but looks like the might spin it out) have so many patents that it's probably mutually assured destruction and a huge gamble if either sues the other. I think they just prefer having the tools proprietary, not just to lock the customers in (though they like that) but also to keep away from the hairy corners (avoid defects in the hardware design that could produce bad results or fry the chips).
The bitstream format is an obstacle but can be reversed. It's already been done for the Xilinx 7 series, lattice ecp5, and others.
However, that does NOT solve the actual main problem - timing models.
Timing models are huge databases hundreds of megabytes for a single FPGA that provides exact routing delays and propagation delays for groups of functional gates in the fabric. They are developed over months of painstaking analysis, debugging and tooling by the vendor.
Timing models are what let's you say "please make this IP run at 166mhz" and the fitter knows exactly how hard to work placing and connecting LUTs so that is possible.
Then, the timing analyzer will check the maximum frequency of that clock domain and ensure, using the timing models, that the specified frequency can be reached at all 4 process corners (PVT).
A typical design will have usually anywhere from 5 to 30 clock domains.
So if you have no support for the timing model, you effectively are not able to ever optimize your fitting process, and you have no idea if your FPGA will have its state machines explode when it gets a bit warmer than ambient.