Intel Gears Up for FPGA Push
nextplatform.com
nextplatform.com
The ability for anyone to develop software for CPUs at zero-cost is an amazing freedom. You literally cannot do that on certain fpgas--not only do they cost thousands, so do the tools to actually create a working design.
Until that changes, FPGAs will always be niche.
That being said, there are soft-processors that have been written in VHDL/Verilog and can be incorporated into a FPGA design, and those CPUs do often have GCC support (Microblaze Xilinx, NIOS 2 Altera, plenty of MIPS CPUs out there, etc.).
And of course if you change supplier you have to relearn everything or pay for an other third-party closed source solution to abstract some of the differences between the environments of Xilinx, Altera/IntelFPGA and friends.
So you got terrible tooling and then a lock in into that tooling.
Not very appealing. Why the f--- would anyone use FPGA's unless they're the only option?
I used to work for a large defense contractor doing FPGA work, and we really went out of our way to try to stick to 100% VHDL to attempt to avoid vendor lock-in. This meant things like writing block ram HDL in a way that it would be inferred by the synthesis tool to use the block ram (and not synthesize out of a bunch of LUTs).
It was a constant struggle to keep things synthesizing correctly. We mostly standardized on using Synopsys Synplify Pro for vendor-neutral synthesis and then using the Altera/Xilinx backend tools for place and route, etc. But even then, we regularly had to whitelist certain good versions of Synopsis' tools.
It was possible to do this with the vendor-specific synthesis tools too, but it was still a massive effort to check that primitives synthesized into the correct elements across different parts/vendors. Ugh, do not miss it.
That would be more than neat, revolutionary really. But it would require buy-in from the major vendors, in the form of publishing many low-level device details.
The compound problem of this in reality is that finding the right people who have the right strain of semi-insanity to do this really well, is very slim. And most of them are tied up with a massive salary from some aerospace company such as Raytheon, GE, Honeywell, or someone like Philips (building medical devices).
Add to that compile times of an hour or a day, and you have a glacial development cycle on your hands.
The gist of which is you can now actually uncheck products you don't need to install, preventing them from being installed.
1. The optimization problems that EDA tools have to solve are about as hard as it gets - NP-hard and sigma 2p problems at a massive scale. I'd argue that these are among the hardest optimization problems in modern computer science today.
2. The number of people with the CS + EE training to develop this software is decreasing. Not a lot of people are getting PhDs in EDA anymore, because:
3. Regular software houses (Google, Amazon, etc.) have about 2-3x the comp as EDA companies. Trust me I've seen a lot of incredibly smart EDA people jump to Google and more than double their salary overnight.
I doubt Intel is going to shift the balance here. It's not easy to develop a fast and efficient SP&R flow.
I wish I had the time to play with it!
Tooling is 100% the thing keeping FPGAs back.
The big consumers of FPGAs don't care if the tooling is terrible, because they're coming from 80s development environments and practices and can afford to dedicate 50-person teams to the job function.
Smaller companies won't touch it because it's terrible, and the vendors won't improve it because there's no demonstrated market for improvements.
Or maybe the market economics work out, but are small relatively compared to the big players, so it moves glacially.
Add a million LUTs and several thousand special-purpose blocks (ALUs, CAM, SRAM, DSPs, etc.) on CPU die that can be reconfigured within a few 10k cycles (ie process context), and then future AI-enabled optimizing compilers could incrementally profile and accelerate applications with reconfigurable resources in new, creative and interesting ways. IRAM and similar zillion core approaches are another approach to solve traditional performance locality bottlenecks: distribute processing, interconnect and RAM amongst each other. It makes the most sense to co-evolve such a radical undertaking gradually with compiler vendors and large customers so that support and usability is working from launch, rather than just throwing silicon over the wall and using the "hail mary" method of product (non)design.
PS: Intel could push almost any new on-die CPU technology service/product into the mainstream right now on the server side because of their de-facto monopoly oligopoly dominance. I don't think they will as there is immense organizational pressure to not innovate too much.
LUTs are implemented as a cascade of muxes that are controlled by local latches that are set at configuration time through a scan chain. Scan chains are key to letting you configure all these LUT control latches across the whole chip in an area-efficient manner, but there's a severe time vs area trade-off since you're essentially configuring one LUT control bit (a vanilla 6-LUT has 2^6 = 64 control bits) per scan chain cycle per scan chain. FPGAs already suffer ~30x area overhead for LUT logic compared to equivalent ASICs, so adding a lot of additional area to enable much faster reconfiguration seems like a tough sell. Matrix addressing as used in DRAMs and LED/LCD displays is another sweet spot on the time vs area trade-off curve. Not sure if anyone's ever done an FPGA using that as the configuration topology, but it'd probably make a fun dissertation topic.
It will be interesting to see if Intel replaces the ARM CPU with a variant of x86.
I would expect a Xeon+FPGA to be at a very different price point from the ARM+FPGA devices, it will be interesting to see whether Intel carry on selling ARM ones or also produce an Atom+FPGA equivalent.
Electrical and computer engineering individuals want cheap big FPGAs too! (Speaking for myself of course.)
Or just release a cheap pcie card! Cmon guys! Oh wait you've forgotten how important it is to get a developer community started around your hardware cause it's been so long since the x86 was new!
Intel: Huh? Developers? But they don't buy servers full of chips. Why do I have to sell to developers?
Both would have significantly more traction if they offered a reasonable desktop class machine, but they don't seem to be able to do it. For some reason there are dozens of RPi type devices but making a $200-$300 device with a reasonable set of expansion ports (sata+USB3+pcie gen3+m.2+10g ethernet+etc) seems all but impossible. POWER has a similar problem, with the only inexpensive devices being old NXP parts (yah NXP is still selling a core that turns 20 years old this year https://www.nxp.com/products/microcontrollers-and-processors... thats great if your still selling car parts from 20 years ago, not so great if your designing a part today.).
So, a company may only be making a dollar on each machine sold, but is gaining thousands of hours of beta testing, or actual engineering in return.
I should add though, that this months basement hacker is next years businessman buying in lots of 5k to fulfill some small niche market. Do that enough times and eventually the next netflix/facebook app just happens to be running on your hardware.
1) Make your own SoC designed for the purpose. You can tailor it to meet your requirements precisely, but given the low volumes you'll be selling, your system will be at least $10,000 a box, likely more.
2) Use a designed-for-mobile SoC. This will hit your $200-$300 price point pretty easily (perhaps even sailing under it by a big margin), but the IO options will be bad, because mobile SoCs don't neeed SATA or PCIe. CPU performance is likely to be underwhelming because mobile SoCs are designed to hit a power consumption level that won't drain mobile batteries or make mobile devices overheat. Almost all the cheap devboards you can get today are in this category.
3) More recently we're starting to see server or networking SoCs, which gives you an option a bit like 2 but with different cost and expansion tradeoffs. Generally a bit more expensive than 2 because networking won't hit mobile volumes, but better I/O capability. CPU power may still be less than you might like. Examples here are the Macchiatobin board and the dev box that Socionext announced last week.
I don't think anybody disagrees that a proper desktop class machine for these architectures would be great; but there are huge economic barriers to getting there. Personally if I was looking at getting a new ARM setup I'd try something in class 2, likely the Socionext box when it becomes available (end of the year, I think they said).
Broadcom designed a custom SoC for the Raspberry Pi 3 and everything still runs over USB 2.0.
[1]:https://secure.raptorcs.com/content/CP9M01/purchase.html
That way instead of spending millions designing a SOC for each tier of devices you design a "generic" device and sell the ones with busted cores/cache/ram channels/whatever as low end developer machines after fusing the broken functionality off.
IMO the problem with FPGA or even silicon dev is mostly tooling. I often say this, the biggest contribution to Open SOurce movement is not of linus/linux ..its gcc. just imagine if 'they' could tangle up every bit of open source code in IP litigation emerging from proprietary compilers.
What long list of applications can CPU+FPGA bring us that a CPU or GPU can't?
The optimization needs to beat out the inefficiency of the FPGA compared to an ASIC though (not to speak of development time). But sometimes, you can do just that.
For more modern hardware software implementations can't reach absolute accuracy in real time but FPGAs big and fast enough to handle these designs are prohibitively expensive (if they even exist in the first place). That's why FPGA GameBoy emulators are not popular and you probably won't see a FPGA PS3 emulator any time soon (unless the pricing changes dramatically in the near future).
FPGAs shine when you need to process very high throughput data with low (or at least constant) latency or for special-purpose algorithm with no hardware acceleration available on CPU or GPU (video codecs, crypto, computer vision etc...). But in general when an algorithm becomes popular enough (AES, SHA-256, H264...) hardware support is backed into the ASIC eventually with much better performance and power consumption than a FPGA.
I see a potential for FPGAs in CPUs for professional applications if you want to process big datastream without having to rely on external hardware. For instance I work in broadcast video transmission where we routinely handle uncompressed HD or even 4K streams, being able to prototype directly on my workstation's CPU would be pretty cool.
As for general purpose programing language optimization I have a hard time imagining what it would look like. The problem is that you don't code for a FPGA the way you code for a CPU. To put it very simply on the CPU serial is cheap while forking and synchronizing threads is expensive. On an FPGA it's basically the other way around. Automatically transforming one form into the other automatically in a compiler or JIT sounds very much non-trivial. Maybe I just don't have enough imagination.
With commercial FPGAs you have a problem that you can't generate the bits yourself but have to use the vendor tools. These are pretty simplistic and will take minutes or hours to compile non trivial circuits. What I mean by this is that they flatten the netlist at some point so compilation time grows non linearly with circuit size. So if your design has five identical blocks, adding a sixth might make it take twice as long to compile.
A second problem is that it takes a while to load the bits into the FPGA. The old XC6000 could be accessed as RAM and quickly rewritten but all other models serially shift in the configuration bits. Xilinx allows partial reconfiguration where you can load the bits to one part of the FPGA while the rest continues to work but that is pretty hard to use and the competitors don't even offer this option.
I'm sure I missed something in your first paragraph though, because the process that you described sounds very much like a normal modern JIT compiler. In V8 and SpiderMonkey for example, the engine starts with the first tier, which is a baseline interpreter that collects type information on each function call, and then if a function is invoked enough times, it compiles the function to machine code using the type information collected in the first tier. There are more tiers up the chain, slower to compile and faster to run, where dataflow analysis is done on the hottest functions, which get optimisations such as inline caching, etc.
All that is done on the CPU of course, so I'm curious how an FPGA could help in this process?
FPGAs don't help JITs at all. I was saying the opposite: if certain problems (slow compilation and slow reconfiguration) could be solved then JITs could offer the opportunity to generate hardware acceleration on demand that would be tuned for a specific execution of an application.
What is the price of the new FPGA add-on cards? That is what will make the difference here for a 'new push'.
Most complaints about vendor software are closer to the front of the flow. Examples...
* Poor language support (applies for both SystemVerilog and VHDL)
* Not enough transparency and access to primitives and IP blocks, resulting in poor ability to automate
* Generally buggy elaboration and synthesis results, sometimes even causing the tool to crash
My opinion is that the FPGA companies spend too much money improving the HLS (c-to-gates) and IP wizard experience, in an attempt to make their devices more accessible to the mythical software engineer who wants to use an FPGA.
They should have spent that time and money supporting language standards, and improving the RTL experience, which is how most engineers use their products.
In the unforseeable future, then who knows, maybe the predictors will become so flexible that even conventional CPUs evolve into something like basically FPGAs.
I've had a vision for "the future of computing". FPGAs that reconfigure themselves (<-these exist) to become whatever macro-level hardware assets your computer needs. Running tons of SHA-256 encryptions per second? The CPU/OS/OpenSSL (or whatever) detects this condition and switches from hand-coded ones, to an IP core that comes with the CPU. The CPU flashes the FPGA's to become SHA-256 "CoreS" and now you're running 4096x the output with less heat. (As CPUs are designed for one thing "few, large, complex cases" while FPGAs are perfect for many, parallel simple cases" even more than a GPU.) Now, you shutdown your encoding and switch to video encoding, or Doom 2019 and your CPU reflashes (Alteria specialized in PARTIALLY flashable FPGAs so you don't have to nuke the entire FPGA, only sections) and adds cores for video, or physics, or "shader units".
This would be hard for a single person, but any large company could handle making this. You can even do it with off-the-shelf FPGAs. The biggest problems are 1) bandwidth. The "macro" function size needs to be bigger than the latency hit you take for asking the FPGA over computing it internally. (Intel's on-CPU FPGA would be insanely fast access.) and the other one is 2) How do you get people to use it! The simplist is of course, only supporting people who actually request it. But, you can take libraries that perform common, encapsulated macro functionality, like OpenGL, a physics library, or OpenSSL, where people don't care about the inner code ("how it gets done") but instead the result. Asking for floats to be multiplied would be bad. But asking for a cross-product would be much faster. Asking for an SHA-512 key would be super.
And the benefit here is, you don't have to hardcode that functionality into a CPU. The FPGA can have NEW or improved IP cores downloaded with Windows Update every week.
Back when I was in college, I actually bought a Lattice dev kit with a PCI Express card, dual-gigabit ethernet, DDR3, and an near top-of-the-line FPGA on the board and it cost a mere $100. Unfortunately, I was a more software guy so I really got in over my head (plus health issues set me back and have never let up since), so I never got a working prototype built.
But it's still there! A huge opportunity waiting to be seized that could really become another tier of "the standard PC." In the same way we think of SSDs as "almost RAM" scratchpads, or GPUs as "CPUs for massive amounts of simple decisions." Well an FPGA is the "GPU of GPUs". Even simpler decisions and insanely fast and parallel even at "low" (by CPU standard) clockrates of 400 MHz.
Here's an older (2009) project/research article that inspired me called the 512 FPGA cube.
http://cc.doc.ic.ac.uk/projects/prj_cube/Welcome.html
http://cc.doc.ic.ac.uk/projects/prj_cube/spl09cube.pdf
And here's a direct link to the data table results between FPGA, FPGA cube, and Xeon (and cluster of Xeons) trying to do the same work:
https://i.imgur.com/byjmEDG.png
Those are massive differences in numbers in both power efficiency and compute rate. 72,000 Watts of Xeons to get the speed of a single 832 watt cube. That's two orders-of-a-magnitude!
I mean, imagine a world if they bothered to make FPGA's you could plug into Ethernet, and make them configure themselves according to a simple programming tool that was "easy" for normal programmers to exploit instead of requiring intense understanding of logic gates, propagation delay, and so on. A tool that wasn't "as fast as" a dedicated engineer, but 90% (or even 70%) as fast at zero cost and effort. All a sudden you could run tons of programs and macro-sized functions like you had personally stamped them into a printed circuit yourself, but without spending millions on development.
I'm honestly not sure why this hasn't already happened. I can't be the only "smart" person who came up with this idea. And the research (and practice with bitcoin miners) all points to a huge opportunity to be exploited if they could lower the knowledge barrier-to-entry so you can basically "push a button" and unleash an FPGA at a problem. Imagine LAPACK and BLAS with FPGA support.
So what can FPGA do? Fast, low latency, high bandwidth interaction with peripherals. The irony here is that to have this work out, you kind of want to have your peripheral connected to the FPGA.. which takes away all the fun from the reconfigurable stuff, because you can't reroute your PCB. So now 99% of FPGAs deployed end up running in the same configuration always and companies with the necessary scales pour it into ASICs.
FPGAs solve a niche problem of interacting with very fast, massively parallel data buses and systems (think CCD sensors, ADC sampling, ..) that a linear execution, Turing style processor isn't suitable for. And pretty much only for applications where you don't have the volume to convince a chip manufacturer to put your peripheral into silicon.
perhaps you might find it interesting. have fun :)
0) Trying to do automatic parallelization is something we've been working on for 50 years and we still haven't solved in any practical degree. You can't just slap a #pragma on C/C++ code at this point to say "run this on some non-CPU architecture" and expect to get good performance.
<offtopic>Hmm, where have I heard something like this before ... ah, yes, the brain - CPUs/FPGAs are like reason and instinct, because reason deals with "few, large, complex cases" and instinct has "many, parallel simple cases". The brain has its own CPU/FPGA divide.</>
FPGAs are quite good at parallel simple cases, that is correct, but they would lose to GPUs in performance/watt in most cases. Where FPGAs really shine is in parallel complex, non-uniform cases, especially cases that don’t map well to the classic CPU instructions, but can easily be performed with small latency on FPGAs.
All that said, these are golden years to be a low-level programmer who understands parallel algorithms whether you work in Tech or at a hedge fund because there just aren't that many of us.
But the real problem with FPGAs is that even if they find another lucrative application where they excel relative to GPUs, Nvidia can simply dedicate transistors in their next GPU family to erasing that advantage as they did with 8-bit and 16-bit MAD instructions in Pascal and with the tensor cores in Volta. Too bad they don't care about latency or I believe they could disrupt FPGAs from HFT in a year or two when someone started using them and started winning.
Also, you'd have to know what you'd need... before you need it. Which is kind of impossible. By time you know you need tons of integer units, you probably could have started working on them. That is, if you need to rapidly switch, then your workload is pretty rapidly completed to begin with.
However, I don't thing they need to reconfigure that fast. Once every second would be enough to keep up with most workloads. Most "heavy duty" workloads aren't changing that rapidly. You load a video game, it's a videogame for hours you play it. You load a web server, you're going to be doing SSL.
If you need much more fine control, it'd probably be better to treat the problem at a much higher level ("I need more SSL keys / sec", instead of "I need more integer adds to make SSL keys") or add another FPGA (one for each use case or set of use cases, ala one for web server keys, one for deciding some other major web server feature).
Of course, I'm no expert in the field. I'm just a guy with an idea and some experience / research into FPGA's as reconfigurable logic units.
Frankly Intel has done an adequate to good job of providing open source software to interact with all of their hardware for as long as I've been computering. They're the one vendor that has never told me to install a proprietary driver.
That said, you can't escape that with their nearest competitor, so it's a bit of a moot point for now.
I'm not really seeing loads of people taking advantage of the feature however. The platform is cheap, the technology is available but its just way too weird an architecture to become mainstream.
There are numerous benefits: the CPU can create a linked list or graph, and the memory will still be valid on the GPU. CPU / GPU atomics are unified, and GPUs can even call CPU functions under AMD's HSA platform.
* https://images.anandtech.com/doci/7677/20%20-%20HSA%20Use%20...
* http://developer.amd.com/wordpress/media/2012/10/hsa10.pdf
* https://www.anandtech.com/show/7677/amd-kaveri-review-a8-760...
---------
I think Intel had a similar technology implemented on their "Crystalwell" chips, which were basically an L4 cache which provided a high-bandwidth link between the CPU and GPU (although not quite as flexible).
No, its not an FPGA, but OpenCL / GPGPU compute seems to be a bit more mainstream than FPGA compute at the moment. I haven't seen too much excitement in general for this feature however.
AMD sells a consumer product. For most consumers even a smartphone offers enough CPU and GPU performance. The content producers who care about performance usually buy the best CPU and GPU. HSA isn't available on AMD's Ryzen or Threadripper processors.
Intel is trying to sell to datacenters where performance or energy efficiency is a major selling point.
Raven Ridge will be based on Zen CPU cores and Vega GPU cores. But naturally, Raven Ridge will be slower than Threadripper because the GPU will take up some space (that otherwise would have been additional CPU cores).
Rumored specs of Raven Ridge APU is 4 CPU cores and 11 GPU Compute Units. In contrast, Threadripper is around 16 CPU Cores and Vega 64 is 64 GPU cores, separated by a PCIe x16.
So basically, its the price you pay for sticking so many things onto a single package. There are thermal limits, as well as manufacturing limits (ie: practical yield sizes) to how large these chips can be.
If you want the best of both worlds, like an EPYC CPU with Vega 64 or a high-end NVidia Pascal / Volta chip, you'll need to buy a dedicated GPU and a dedicated CPU. True, a hybrid chip like Raven Ridge (or any of the AMD HSA stuff) has benefits with regards to communication, but the penalty to CPU speed and/or GPU speed seems to be huge.
-----------
I personally expect that if any "mainstream" FPGA solutions come out, they'll be connected to the PCIe and not merged into the CPU. There seems to be just too many heat and manufacturing issues to make a merged product compared to the standard PCIe x16, which is quite fast.
Alternatively, certain tasks (like Cryptography) can be accelerated using dedicated instructions, like the Intel AES-NI instruction set. Or Intel's Quicksync H.264 encoding solution. Fully Dedicated chipspace (like AES-NI) is way faster and more power efficient than FPGAs after all.
If there turns out to be a market (which comes down to tools and cloud platform to make this existing tech much more accessible) then it'll be huge and they'll need to be part of it.