This could change, and I'd love to see that, but without improved "programmability" FPGAs will remain niche. I haven't followed the latest developments very closely, but Xilinx seems to be moving in this direction with Everest through the addition of vector cores.
As a disclaimer, I've never personally worked with FPGAs, but a lot of my coworkers have, and I'm mostly parroting their views. And those views are...
FPGAs are a dead-end technology. For a while, they were considered the "obvious" next frontier of HPC, just once we got the programming model sorted out. And they stalled at that phase for over a decade, until CUDA came out and everyone realized just how much better GPGPUs were at providing the necessary HPC speedup while requiring far less development time.
So how do you fix FPGAs? Well, you start substituting actual hardware logic instead of emulating everything with LUTs. And as you do this, you start to end up at coarse-grained reconfigurable arrays instead.
Source: Professional FPGA designer.
A multiplier is a multiplier is a multiplier. You can spin your own, but it is so common that every FPGA vendor says, "If you need to multiply, you can use this block." You want to write code that is general:
A <= B * C + D;
And have your synthesizer go, "This is a multiply accumulate! I can fit that into the following blocks: Look up tables or a DSP slice. I'll use the DSP slice - it is faster and smaller." Note that you didn't directly call the DSP slice, you just said "multiply". It doesn't always work this way, but that's the goal.Not every FPGA has a dual core cpu - the Virtex-5's and 6's started that trend with a PPC block, and it really hit its stride with the Zynq-7000 (when Xilinx switched to ARM cores and brought the price down significantly). Again, if you're doing things a CPU can manage, you CAN do them with logic, but why not use the embedded core?
The addition of the ARM cores was a brilliant move by Xilinx because a number of embedded systems out there used FPGAs for fast response, high speed interfaces, high speed datapaths, and as glue, but many included a separate processor to handle "housekeeping" tasks.
Xilinx noticed this, and said, "If you want, you can choose the chip that has a processor core in one corner instead of reprogrammable fabric there." Bam, tons of sales - because now instead of two chips, I need one. Integration.
They're furthering their exploration of that interface with HLS and their newer toolsets. The idea is to make the algorithmic division between the two things null - you can seamlessly switch between control and data dominated computation.
But let me take one step back.
What's the difference between a state machine and a counter?
Nothing. They're a cloud of 'next state' logic, a current state, and an output based on the current state and/or the current state + the current inputs (Mealy/Moore/Medvedev). The 'next state' logic happens to also follow the rules of arithmetic, which is what you're interested in when you're using it as a counter, obviously. But it's a state machine.
So what's the difference between a CPU and a state machine?
Again...nothing. A CPU is just a complicated state machine (or a set of interacting state machines).
To dislike hard macros which are fast and common is to not grok FPGAs. To say you work with FPGAs professionally but you consider them a dead end, when they're built out of the same logic and blocks that underlies everything that's ever been done with a digital computer is...confused, at best.
We're one step "above" transistors in the abstraction hierarchy. But those transistors would implement things like counters (assuming you're not building an analog computer!), and you'd be right back to an FPGA (well, an ASIC, so YOU get to choose what blocks get included or not!).
A lot of people who design ASICs do it from HDL's - they even use synthesizers! But the synthesizers are targeting a library of parts for a silicon process. The only difference is that the blocks being targeted on an FPGA already exist - you can't move a block over to get better timing like you can with an ASIC, or widen transistor ratios to drive more current, etc - you have to use what's on them.
So...yeah. FPGAs use the fundamentals of all computing directly, they're not a dead end unless we all decide to switch back to analog computations or quantum computers or something, and the hard IP is really important.
Even if spinning an ASIC decreases in price to a few grand (crazy hypothetical), they'll be programmed like you programmed an FPGA - full of the same primordial soup components right above transistors that does all the work in the digital abstraction.
FPGAs are a dead end, but your comment didn't even address for what goal they are a dead end. If you'd take a look a look at a context of this thread, you'd realized we're talking about general applications.
---
> To dislike hard macros which are fast and common is to not grok FPGAs. To say you work with FPGAs professionally but you consider them a dead end, when they're built out of the same logic and blocks that underlies everything that's ever been done with a digital computer is...confused, at best.
Or maybe it's understanding the limitations of the technology you work with. I also never mentioned or implied disliking hard IP.
> What's the difference between a state machine and a counter?
> So what's the difference between a CPU and a state machine?
What's the difference between the universe and a state machine? What can be encoded as a state machine is largely irrelevant. I can make a CPU inside of minecraft but that's not very useful.
> FPGAs use the fundamentals of all computing directly
This is simply not true. FPGAs emulate the "fundamentals of all computing". The emulation is not efficient, hence why hard logic is so important.
Unroll your calls, assign a bit to each statement of your program, create state machine that compute state bit transitions and resource changes. And off you go!
You get TTA that is as flexible as can be.
The optimal RTL is slightly more complex, but also does not require hand written RTL.
In software you allocate and deallocate memory, which is your dynamic resource. You don't have that in HW. You can implement a memory manager in HW, but your system will not be flexible enough to be soft. Also you can synthesize your CPU on FPGA, but the performance is poor compared to a hard CPU on FPGA.
Currently FPGAs are becoming popular in accelerating software functions to offload the CPU. The most powerful FPGAs have a lot of fixed logic CPUs in them (i.e. hard macros).
You can't just say that I'll just go ahead and run my software functions as hardware. For certain type of workloads you always need software, in practice (not in theory).
In a shared-memory system, just implementing libc (or the Java VM, or the Erlang VM) on an FPGA might be a win. It has to be enough bang for the buck for FPGA, but not so much that somebody would make fixed-hardware for it.
On that note, haven't networking end-points had fixed hardware also for ages now? Maybe it is inevitable that a successful application of FPGA's breeds interest in fixed-function hardware for it.
And was it a pleasurable experience?
Do you really think people will rewrite all their software for fpgas?
And what if you're memory or I/O bottlenecked as so many processors are today?
> Do you really think people will rewrite all their software for fpgas?
A compiler will do it for them. If technology moves in that direction they don't have any choice anyways.
> And what if you're memory or I/O bottlenecked as so many processors are today?
I/O will never run as fast as the CPU.
In the era of kernal bypass and modern IO accelerators seen in some data centers today, I wonder if this is even true today?