Low-latency interaction with hardware. In this case, it was possible to start outputting a TCP reply before the end of the packet it was replying to.
In general, FPGAs win for integer and boolean functions which are amenable to deep pipelining, or are capable of very high parallelism. One of my colleagues produced a neat chart of what level of parallelism was best suited to which of CPU/GPU/FPGA.
FPGAs don't have an advantage if you're memory or IO bound, and they have terrible power consumption.
First, of course, are very efficient Ethernet ports and many other efficient hard blocks.
Second, you can fuse many operations into one. For example, you cannot fuse 3 FP adds into one operation nor on GPU neither on CPU. It is often 2 or 4, rarely something inbetween or outside of those. GPU and CPU operations on vectors can be wasteful as vectors can be underutilized and on FPGA you can create circuitry that fits problem rather well.
I also think that you are not quite right about I/O or memory boundness. In FPGA you can add another I/O controller and use extra device for I/O, loading spare resource in FPGA. Same for memory - DDR controller synthesized into cells won't be very optimal, but you can have many of them nevertheless.
Having two DDR controllers helps the overall memory bandwidth number if your reads are not localised. If you're doing a lot of data-dependent loads this doesn't help at all (e.g. scrypt).
In all cases it's much harder for software developers to develop for FPGA, so this cost needs to be factored in. They're very good in their niche, not a general purpose silver bullet.
Your example with scrypt also does not access much of RAM, especially with Salsa20/8 standard function. As far as I can see, it also has parallelization parameter, and computation within top loop can be done in parallel.
Yes, it is hard to program for FPGA. But not that much - I myself programmed a system that performed translation from (pretty much high level) imperative description of algoithm to synthesable Verilog/VHDL code. In a one and half of month.
In my opinion, programming for FPGA is very entertaining, especially if you do not write V* code by hand.
Here's the thing. FPGA performance has nothing to do with it. They do what CPUs can't.
You don't choose to plug an FPGA in place of a CPU, you plug it where you can't.
Tying peripherals together, glue logic, bus connection. Some FPGAs have a builtin CPU (or you can plug a soft one together with your circuit)
Maybe they will go for in-the-fly reconfiguration for specific computations (as: load your specialized circuit in an FPGA and fire away)
I read an interesting quip somewhere on software/hardware development: 'civil engineering would look very different if the properties of concrete changed every 4 years.'
If at some point we stop scaling chip performance. And many-core-integration in/on a single chip/die stops making sense. Then glue logic starts to look like a key differentiator. And control over glue logic starts to look like control over profits.
Intel ate the chipset for performance reasons and so they could shape their own destiny.
If there aren't fundamental breakthroughs to preserve performance scaling as we know it, then I see this as more of the same.
Of course, it depends entirely on what you're doing with them. Keywords: horse, course, different.
And that's where FPGAs shine; that small to medium volume market where small companies are doing innovative things but don't have the millions required to risk building an ASIC.
I don't disagree with you on that count. Especially in this case (since for most cases a processor does a job with a better cost/benefit), FPGAs shine on very specialized/heavy computation tasks.
I always wondered, given that Intel's processors already have a pretty large gap between their instruction set and their real microcode, whether it would make sense to have a nominal "CPU" that, when fed an instruction stream, executes it normally on general-purpose cores, but also runs a tracing+profiling JIT over it to eventually generate a VHDL gate-equivalent program to jam into an on-core FPGA. "Hardware JIT", basically, with no explicit programming step needed.
http://www.linusakesson.net/scene/parallelogram/
"For this demo, I made my own CPU, ... cache, ... blitter with pixel shader support, a VGA generator, and an FM synthesizer."
In his explanation for why he wrote his own CPU in the FPGA, Linus explained "...I was able to take advantage of the added flexibility. For instance, at one point the demo was slightly larger than 16 KB, but I could fix this by adding some new instructions and a new addressing mode in order to make the code compress better."
Also, dude has a Symbolics Space Cadet keyboard. Respect.
I heard counter-arguments to the tune of 'hardly anyone wants to program their FPGA' which sounf strange to me: after all, hardly anyone wants to program their pixel shaders, either.
When it comes to performance, you should look at FPGA in terms of performance per Watt. Generally speaking they outperform GPUs by an order of magnitude in FLOP/W*s [2], which in turn already have ~ 3x-5x advantage over Xeons [3]. This measure is the most important one when it comes to the question, how many chips you can put in a given rack. FPGAs are still held back in terms of cost per dollar invested, since they have been quite pricey - with Intel this could change.
[1] http://stackoverflow.com/questions/17256040/how-fast-is-stat...
[2] http://synergy.cs.vt.edu/pubs/papers/adhinarayanan-channeliz...
[3] http://streamcomputing.eu/blog/2012-08-27/processors-that-ca...
Why would it? The price of an FPGA is not really determined by its production cost – for medium-size Xilinx FPGAs for example, the ratio of price/chip production cost is on the order of 50.
The paper you link to is measuring MSPS/W (mega samples per second) and the algorithm they are studying relies on fixed point. It uses built in DSP blocks in the FPGA that are integer only. There is no floating point so it is incorrect to say this shows FPGAs give better FLOPS/W. It isn't all that surprising the FPGAs are doing better, the GPUs are all about floating point which isn't being used here.
Their GPU implementations use floating point as well as int and short. The efficiency barely differs between them showing that this particular GPU wasn't optimising with integer power efficiency in mind (which an FPGA implementation relying on DSP48s very much is).
Any example of a GPU with "hundreds or thousands of microprocessors"? Nvidia Titan X has 12 [1] microprocessors by your definition.
[1]: SM, Streaming Multiprocessor in Nvidia's terminology. Smallest unit that can branch, decode instructions, etc.
An AMD Radeon R9 290X has 2816 stream processors (44 compute units of 64 stream processors) per their terminology. There is only 1 instruction decoder per compute unit, so a stream processor cannot completely branch off independently, but it can still follow a unique code path via branch predication. This is kind of comparable to an Nvidia GPU having "44 streaming multiprocessors".
But whether you call this 44 or 2816 processors is irrelevant to my main point: a processor that has to decode/execute 44 or 2816 instructions in a single cycle while supporting complex features like caching, branching, etc, is going to be less efficient than a FPGA with hard-wired logic (edit: "hard-wired" from the view point of "once the logic has been configured").
gchadwick also said integer workloads were "not power efficient" on GPUs, but that's also false. Most SP floating point and integer instructions on GPUs are optimized to execute in 1 clock cycle, so they are equally optimized. And of course integer logic needs fewer transistors than floating point logic, so an integer operation is going to consume less power than the corresponding floating point operation.
https://www.altera.com/content/dam/altera-www/global/en_US/p...
FPGAs have high speed links and can perform thousands of operations in parallel. When you have a huge amount of data to process in real time, an FPGA (or ASIC) is often your only choice.
Video works really well because its discrete (60fps) and an FPGA just can't hit the clock speeds an asic will. If you already have a semantic clock requirement that is low, streaming works great.
FPGAs are essentially re-programmable hardware, so they tend to outperform CPUs/GPUs when you program them for a specific task. They don't have to deal with most of the overhead that the more generalized platforms deal with which is why they dominate in the small input sizes. However, with FPGAs you're trading space (silicon) for that re-programmability so you can't have as much hardware in the same area as say a GPU. Thus, when the data sizes have saturated the available hardware of the FPGA for computation, the GPU begins to outperform. Due to the decreasing node sizes (28nm, 22nm, etc), we can fit more programmable logic into the same area, which causes the chart I mentioned above to shift more into the FPGA's favor.
[1]: http://www.researchgate.net/profile/Sam_Skalicky/publication...
But yeah pretty much any repetitive computation. 3DES enc/decryption is another example (although I think people normally use ASICs for that when they want to do loads of it).
FPGAs are like GPUs with no floating-point, no caches, and limited local memory. But if you can implement a kernel in FPGA with comparable memory bandwidth, you'll usually outperform GPGPU while using as little as 1/50 the power.
FPGAs have several orders of magnitude lower latency than GPGPU. GPUs have memory access latency of 1 microsecond, getting something useful out of them >1 ms. FPGAs can have state machines running at 200 MHz, or 5 ns cycle time.
> FPGAs are like GPUs with no floating-point, no caches, and limited local memory. But if you can implement a kernel in FPGA with comparable memory bandwidth, you'll usually outperform GPGPU while using as little as 1/50 the power.
Some FPGAs do have floating point hard blocks. Integrated SRAMs (syncram) can be used as caches and usually are. FPGAs usually have DRAM controllers as hard blocks, so local memory is not that limited. Unless you consider up to 8 GB (newer models up to 32 GB) limited.
Source: http://en.m.wikipedia.org/wiki/Financial_Information_eXchang...
You tell a CPU what to do. You tell an FPGA what to be.